Skip to content

Make container (and every engine's) step output discoverable from the module - #159

Merged
jgruberf5 merged 3 commits into
stagingfrom
feat/154-container-log-discoverability
Aug 19, 2026
Merged

jgruberf5 merged 3 commits into
stagingfrom
feat/154-container-log-discoverability

Conversation

@jgruberf5

@jgruberf5 jgruberf5 commented Aug 19, 2026 •

Copy link
Copy Markdown
Collaborator

Description

GET /api/project-modules/{id}/logs read DeploymentLog rows — which only the retry path ever writes. Every engine (opentofu, container, ansible, ssh, cli, kubernetes, tmos) streams its step output into Task.logs. So a module that had just applied with ten lines of output got:

200 {"module_id":215,"total_logs":0,"logs":[]}

indistinguishable from "this step produced no output", and it sent the reporter to docker logs on the host. The task id — the real handle, GET /api/tasks/{id} — was not reachable from /status or /deployments either; /deployments returned a deployment id that looked like the log handle but was not. Anything driving Forge headlessly diagnoses a failure through these endpoints, so the empty success actively misleads.

Fixes #154

The issue's first option, but for every engine

The issue frames this as container-specific and offers "return the task logs for container-engine modules." Tracing the writers showed the gap is engine-agnostic — DeploymentLog is vestigial everywhere, and all eight engines go through the same task.logs path — so the fix applies uniformly rather than special-casing one engine.

endpoint change
/logs DeploymentLog rows still win when present. Otherwise serve the newest task's logs with source: "task" and task_id, so the caller knows the provenance and can follow to /api/tasks/{id}. limit honoured as a tail; level becomes a best-effort filter on the engines' own markers. A module with nothing at all returns source: "none" plus a hint naming /api/tasks?module_id= — still 200, but no longer silent.
/status latest_task_id — the live lock holder when a task is running now, else the newest task row.
/deployments task_id per row. Deployment has no task_id column, so create_deployment_record (the one writer all eight engines go through) records it in meta_data. No migration. Pre-existing rows report null.

All existing fields are unchanged — the shapes are backward compatible, and the one pre-existing test asserting total_logs == 0, logs == [] on an empty module still passes.

Why not 404 (the issue's second option)

A 404 from /logs would break any existing client that treats it as "no logs yet" (the pre-existing test does exactly that). Serving the real logs with provenance is strictly more useful and strictly less breaking.


Architectural Decision Record (ADR)

  • N/A (Bug fix or minor docs tweak).

Type of Change

  • New feature (non-breaking change adding functionality)

Additive fields only. None of the three routes declare a response_model, so no schema changes — but the /logs route docstring is emitted as the OpenAPI operation description, so openapi.json and api-generated.ts are regenerated (JSDoc only; no type shapes change). (Corrected after CI caught the stale spec — my original claim that the spec was untouched was wrong.)


Verification & Testing

  • test_routes_project_deployments.py (+5): task-log fallback with source/task_id — the reporter's exact scenario; DeploymentLog still preferred when present; empty module carries the hint; limit tails correctly; /deployments exposes task_id and tolerates pre-Container step logs are reachable only via /api/tasks — the module log endpoint returns an empty 200 and no task id is exposed #154 rows with no meta_data.
  • test_project_module_service.py (+1): latest_task_id is null → newest task → the lock holder when one is running.
  • New test_deployment_record_task_link.py: the shared writer records task_id/celery_task_id.
  • Regression sweep on deployment/module_status/project_module/logs/task: 744 passed, 0 failed. ruff clean (it caught 3 E702 one-liners in my tests before push — fixed).

Environment validation needed

Light — the behaviour is fully covered, but worth one manual check on a live instance since the reporter was on 3.1.6: apply any container module, then GET /api/project-modules/{id}/logs should now return the step output with "source": "task", and /status should carry a latest_task_id that resolves at /api/tasks/{id}.


Checklist

  • My code follows the project's code style and formatting guidelines.
  • I have updated documentation where necessary — the /logs docstring and an inline comment explain the DeploymentLog vs Task.logs split, which is the non-obvious part.
  • I have generated updated openapi types (make openapi-types) — the /logs docstring change lands in the spec as a description.

…om the module

GET /api/project-modules/{id}/logs read DeploymentLog rows, which only the
retry path ever writes. Every engine -- opentofu, container, ansible, ssh,
cli, kubernetes, tmos -- streams its step output into Task.logs. So a
module that had just applied with ten lines of output got

    200 {"module_id":215,"total_logs":0,"logs":[]}

which is indistinguishable from "this step produced no output", and sent
an operator to `docker logs` on the host. The task id -- the real handle,
GET /api/tasks/{id} -- was not reachable from /status or /deployments
either; /deployments returned a deployment `id` that LOOKED like the log
handle but was not. Anything driving Forge headlessly diagnoses a failure
through these endpoints, so the empty success actively misleads.

This is the issue's first option, applied to every engine rather than
container only, because the DeploymentLog gap is engine-agnostic:

- /logs: DeploymentLog rows still win when present. Otherwise serve the
  module's newest task's logs, with `source: "task"` and `task_id` so the
  caller knows where they came from and can follow to /api/tasks/{id}.
  `limit` is honoured as a tail; `level` becomes a best-effort filter on
  the engines' own markers. A module with nothing at all returns
  `source: "none"` plus a hint naming /api/tasks?module_id=, instead of a
  silent empty list. Existing fields are unchanged, so the shape is
  backward compatible.

- /status: `latest_task_id` -- the live lock holder when a task is running
  now, else the newest task row.

- /deployments: `task_id` per row. Deployment has no task_id column, so
  create_deployment_record (the one writer all eight engines go through)
  records it in meta_data; no migration. Pre-existing rows report null.

None of the three routes declare a response_model, so the OpenAPI spec is
unchanged.

Fixes #154

Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4
…g branch

Self-review of #159: the DeploymentLog branch has always returned entries
newest-first (timestamp.desc()); my task-log fallback returned them in
natural (oldest-first) order. A caller treating logs[0] as "most recent"
would get a different answer depending on which source happened to serve
it. Tail first, then reverse, so both sources honour one contract. The
docstring now states the ordering and the three sources explicitly.

Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4
@jgruberf5

Copy link
Copy Markdown
Collaborator Author

Self-review

One defect found and fixed, plus a ruff catch before push.

🟠 Fixed — the two log sources returned opposite orderings

The DeploymentLog branch has always returned entries newest-first (timestamp.desc()). My task-log fallback returned them in natural, oldest-first order. So a caller treating logs[0] as "most recent" — which the existing contract allowed — would get the first line of output from a container module and the latest warning from a retried one, with no way to tell which. That's the kind of inconsistency that looks like "it works" in a demo and bites a script later.

Tail first (lines[-limit:]), then reverse — both sources now honour one contract, and the docstring states it: newest first regardless of source, with the three sources (deployment_log → task → none) listed in preference order. Tests updated to pin it (logs[0] is the last line written; a limit=5 tail returns lines 49..45).

Checked and found correct

  • holding_task_id is the DB row id, not the Celery id — Integer column, and every engine calls module_lock(db, module.id, task_id=task_db_id). So latest_task_id is type-consistent whether it comes from the lock holder or the newest row.
  • Size is bounded: task.logs can be MBs for a tofu run, but the response is capped by the existing limit (1–10000 lines) as a tail, so the fallback can't blow up a client that the DeploymentLog path wouldn't.
  • No response_model on any of the three routes, so the OpenAPI spec and openapi-types-check are untouched — confirmed by grep, not assumed.
  • Frontend doesn't consume /project-modules/{id}/logs (it uses the task WebSocket and bare-metal/k8s log routes), so the ordering fix can't regress a UI; the only consumers are API clients like the reporter.
  • meta_data was unused on Deployment rows from create_deployment_record (it's a nullable JSON column), so writing {task_id, celery_task_id} into it can't collide with anything. Pre-existing rows report null and the test covers that.

Ruff, caught pre-push this time

ruff flagged 3 E702 (db.add(x); db.commit() on one line) in my new tests. Fixed before the first push — the #146 lint failure made that a standing check.

Scope note, restated for the reviewer

The issue is filed against the container engine; the fix is engine-agnostic because the DeploymentLog gap is (only the retry path writes it; all eight engines write Task.logs). That's a wider blast radius than the title suggests, but it's additive fields and a fallback that only fires when the old path returned nothing — the pre-existing empty-module test still passes unchanged.

24/24 in the route file, 744/0 on the broader sweep, ruff clean.

…string

CI's OpenAPI Spec Freshness check failed on #159. The self-review commit
rewrote the /logs route docstring to state the newest-first ordering and
the three sources -- and FastAPI emits route docstrings as the operation
`description`, so that IS spec content. The PR body's claim that the spec
was unchanged because no response_model changed was wrong: docstrings
count too.

Diff is the one description field in openapi.json and the matching JSDoc
in api-generated.ts. No schema or type shape changes.
    make openapi-check -> OK: backend/openapi.json is up to date

Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4
@jgruberf5

Copy link
Copy Markdown
Collaborator Author

CI fix — OpenAPI Spec Freshness

P1 · OpenAPI Spec Freshness failed. The cause was a wrong claim in my own PR body: I said the spec was unchanged because none of the three routes declare a response_model. That reasoning missed that FastAPI emits each route's docstring as the operation description — and my self-review commit rewrote the /logs docstring to state the newest-first ordering and the three sources. Docstrings are spec content.

Diffed the regenerated spec against the committed one to be sure that's all it was:

/paths//api/project-modules/{module_id}/logs/get/description   ← the only changed key

Regenerated backend/openapi.json and frontend-v2/src/types/api-generated.ts; the TS diff is JSDoc only, no type shapes. make openapi-check → OK: backend/openapi.json is up to date. PR body and checklist corrected.

Lesson I'm carrying forward: "no response_model change" ≠ "no spec change" — any route-level edit, including docstrings, needs make openapi-check before push. Adding it to my pre-push list alongside ruff.

@mwiget mwiget left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is right and the fix is the correct shape: DeploymentLog is written only by the retry path, every engine streams into Task.logs, so the empty-200 was structural rather than "no output". Falling back to the newest task and naming the source is better than silently merging the two — a caller can tell which store it got.

Verified against the code, not just the description:

  • Deployment.meta_data (models/project.py:358) has no other writer or reader anywhere in backend/, so stamping {"task_id", "celery_task_id"} there clobbers nothing.
  • Task.module_id exists and is indexed; ProjectModule.holding_task_id is a Task.id (set by EntityLockService with the DB task id from fetch_task_or_raise), so latest_task_id is coherent — preferring the lock holder over the newest row is the right precedence.
  • /status has no response_model, so latest_task_id is genuinely additive.
  • No frontend consumer calls getModuleLogs today (only the API wrapper exists), so "timestamp": null on the task path breaks no UI.
  • Newest-first on both branches is right and worth having stated in the docstring — the two sources would otherwise disagree on logs[0].

Three things below, none of them blocking. Not approving yet only because P3 · Integration Tests · Backend is still running — I'll approve once it's green if nothing changes.

# `id` is the deployment row, NOT the task -- an easy thing to
# mistake for the log handle (#154). Older rows predate the
# meta_data backfill and report null.
"task_id": (dep.meta_data or {}).get("task_id"),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The same id-looks-like-the-log-handle confusion this fixes exists one route down: get_project_deployment_history (GET /project/{project_id}/deployments, the "id": dep.id block near line 303 on staging) builds the same dict shape and doesn't get task_id.

An operator on the project-wide view — arguably the more likely entry point when they don't yet know which module failed — still gets a number that isn't the log handle. Same one-line addition, same (dep.meta_data or {}).get("task_id").

"error": ("ERROR", "✗", "error:", "--- ERROR ---"),
"warning": ("WARN", "WARNING", "⚠"),
"success": ("✓", "SUCCESS", "Complete"),
"info": (),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"info": () means wanted is empty, so if wanted: is false and no filter is applied — ?level=info returns every line, errors and warnings included. That's the opposite of what a caller asking for info-only expects, and it's the one level where the parameter silently means nothing rather than meaning something approximate.

Given free-text logs there's no good positive marker for info, so the honest options are: exclude the error/warning markers instead of matching a positive set, or drop info from the dict and let the docstring say that info is unfiltered on this path. Either beats a silent no-op.

from models import Task as TaskModel

task = (
db.query(TaskModel)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit, for later rather than now: Task.logs is a deferred() column, so task.logs here triggers a second query that pulls the entire log text into memory purely to splitlines() and tail the last limit lines. For a long apply that's a large TEXT round-trip on every poll of this endpoint.

Not worth restructuring in this PR — correctness first — but if /logs ends up polled by the UI, a server-side tail (or at least load_only-ing the columns you need on the task lookup, since you only use id and logs) would be the next step.

jgruberf5 pushed a commit that referenced this pull request Aug 19, 2026
…ywords

Review finding (mwiget): merging #158 as it stood would have closed #94
and #128 -- not because of the regex change, but because the PR's own
description DOCUMENTS closing keywords in backticks, and the parser reads
body text raw. Reproduced against this branch's real parse step with this
PR's body as PR_BODY:

    Parsed closing-keyword issues: ['94', '128', '7']

#94 and #128 are open with their real fixes unmerged in #157/#156; they
would have closed with "Auto-closed by PR #158", wrong issue and wrong PR.
The old parser had the same blind spot (it read ['94'] from this body);
widening the skip tripled the blast radius on a body that talks about the
very forms it now accepts. A parser PR is the right place to close the
class, not the instance.

Fenced blocks are stripped first (they may contain backticks), then inline
spans. Re-ran the full matrix through the real step: every real closing
line in plain text still closes; every example in backticks or a fence no
longer does; a body with both a real "Fixes #94" and a documented
"`Fixes #999`" closes only 94. This PR's own body now yields no issues.

Then ran EVERY open PR's actual body through the patched step -- the
check I should have done the first time:

    #156 -> 128   #157 -> 94   #159 -> 154   #160 -> 99
    #158 -> (none)   #161 -> (none, deliberate Refs #79)   #135 -> (none)

Each PR closes exactly its own issue and nothing else.

Claude-Session: https://claude.ai/code/session_01UpRYiFserdBE5ESHn759N4

@mwiget mwiget left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving — the full pipeline is green on 59670e0 (25 checks, CI Gate SUCCESS), including the P3 integration run I was waiting on, and the head is unchanged from what I reviewed above.

The three points in my earlier review stand as follow-ups rather than objections:

  • task_id is still missing from the project-level deployment history (GET /project/{project_id}/deployments), which has the same id-looks-like-the-log-handle confusion this PR fixes on the module-level route.
  • ?level=info is a silent no-op on the task path — it returns every line rather than info-only.
  • Task.logs is deferred(), so the whole log text loads just to tail limit lines.

None of them are worth holding the fix for; the empty-200 was the real problem and it is properly diagnosed and closed. Happy for the first one to land as a one-liner here or in a follow-up, whichever you prefer.

@jgruberf5
jgruberf5 merged commit 1c4bcc7 into staging Aug 19, 2026
25 checks passed
@jgruberf5
jgruberf5 deleted the feat/154-container-log-discoverability branch August 19, 2026 15:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants