Skip to content

fix(freshness): calibrate prompt by severity so Check Freshness surfaces findings again - #2074

Merged
jung-thomas merged 1 commit into
DEVfrom
worktree-freshness-prompt-calibration
Aug 28, 2026
Merged

jung-thomas merged 1 commit into
DEVfrom
worktree-freshness-prompt-calibration

Conversation

@jung-thomas

Copy link
Copy Markdown
Contributor

Problem

The Check Freshness admin action "returns nothing" — pressing it produces a DONE report with 0 findings even though the LLM runs and consumes tokens.

Root cause (diagnosed on DEV)

The 08-25 release (#2038) bundled a prompt rewrite in srv/lib/freshness-detector.js (commits 8ff849c8 + a3f7eb5c) that shifted the system prompt to "prefer reporting NOTHING over a speculative one." This over-suppressed findings.

Verified with a controlled A/B on DEV (same tutorial, same grounding, same model/deployment):

Prompt cap-self-contained-dev-env
Pre-release (loose) 3 findings (matches the 08-24 report)
Tightened (deployed) 0 findings

Ruled out: the action handler, source markdown (ContentCurrent.sourceContent populated), corpus embeddings (ApiDocs 60/60, Samples 311/311), grounding (12 hits), and the stepRef finding-filter (LLM returned 0 before filtering).

Fix

Recalibrate PRECISION by severity:

  • High/Medium stay strict (confidence-first, prefer omitting speculative ones).
  • Low is now advisory — cosmetic/dated-style staleness surfaces instead of an empty report.
  • All other guardrails (CONTEXT, OUTPUT vs CODE, RESPECT INTENT, SAP CONVENTIONS, SCOPE) unchanged.

Validation (DEV, real grounding, claude-4.6-sonnet)

Tutorial Before After
cap-self-contained-dev-env 0 1 (Low)
hana-trial-advanced-analytics 0 3 (1 Medium, 2 Low)
abap-environment-rap100-early-numbering 0 2 (Low)

The Medium ("JSON_TABLE defines a redundant LOCATION column") is genuinely actionable.

Tests

  • Updated test/unit/freshness-prompt-guard.test.js to assert the new severity-calibrated intent.
  • All 16 freshness unit tests pass.

The 08-25 release (bundled prompt tightening) shifted checkFreshness to
'prefer reporting NOTHING over a speculative one', which suppressed
essentially all findings — the LLM returned 0 raw findings even with
grounding present, so the admin action appeared to 'return nothing'.

Recalibrate PRECISION by severity: keep High/Medium strict (confidence-
first, prefer omitting speculative ones) but treat Low as ADVISORY so
cosmetic/dated-style staleness surfaces instead of an empty report. All
other guardrails (CONTEXT, OUTPUT vs CODE, RESPECT INTENT, SAP
CONVENTIONS, SCOPE) are unchanged.

DEV A/B (real grounding, claude-4.6-sonnet): tutorials that returned 0
under the tightened prompt now return 1-3 findings (mix of Medium + Low),
including a genuinely useful Medium.

Updates the prompt-guard unit test to assert the new severity-calibrated
intent.
@jung-thomas
jung-thomas merged commit 000b1fd into DEV Aug 28, 2026
4 checks passed
@jung-thomas
jung-thomas deleted the worktree-freshness-prompt-calibration branch August 28, 2026 16:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant