Skip to content

Phase 2 — Joule learning-path generator #445

Description

@jung-thomas

Phase 2 — Joule learning-path generator

Sub-issue of #381 (Phase 1 shipped via PRs #401 → #441 + #442). Phase 2 was deferred per the design spec and explicitly scoped out of Phase 1.

What this delivers

A new Joule chat tool that lets developers ask natural-language questions like:

  • "I want to build my first CAP service that uses HANA Cloud and exposes a Fiori UI."
  • "What should I learn after completing the CAP getting-started mission?"
  • "Show me a path from cap-handlers to hana-cloud-deployment."

Joule translates the question into a pathBetween(fromSlug, toSlug) or conceptsForUser(userId) named query, runs it against the HANA KGE named graph, and returns an ordered tutorial sequence respecting prerequisites + the user's existing concept coverage.

Strong demo narrative: "ask Joule for a learning path → see SPARQL → see KGE answer." Both SAP technologies on display.

What's already in place from Phase 1

  • KnowledgeGraphService.pathBetween(fromSlug, toSlug) — declared in srv/knowledge-graph-service.cds as a stub that returns []. The contract is stable; Phase 2 fills in the implementation.
  • KnowledgeGraphService.conceptsForUser(userId) — declared as a stub that returns { learned: [], partial: [] }. Same contract-stable pattern.
  • srv/lib/kg-queries.js exports PATH_BETWEEN_QUERY and CONCEPTS_FOR_USER_QUERY as stub SPARQL strings. Phase 2 replaces them with real queries.
  • srv/lib/kg-sparql-client.js (sparqlQuery) — already wired, with privilege-error / syntax-error / timeout taxonomy.
  • The HANA KGE named graph is populated nightly by the consolidator's graphRebuild step.
  • Joule chat infrastructure exists — see srv/chat-service.cds and the existing getRelevantSteps / checkCode chat tools.

Scope

Tasks

  1. Implement pathBetween SPARQL — a multi-hop query that finds a chain of tutorials connecting fromSlug to toSlug via :teaches → :requires reverse edges. SPARQL property-path syntax (?a kg:teaches/^kg:requires*/kg:teaches ?b) is the canonical shape; verify against HANA KGE's property-path support during implementation.
  2. Implement conceptsForUser SPARQL — joins the user's TaskRecords (completed tutorials) against kg:teaches to compute learned (concept covered ≥1 time) and partial (concept covered by an in-progress tutorial). Also gives the user-coverage signal that improves Phase 1's whatToLearnNext re-ranker.
  3. New Joule chat tool: findLearningPath — registered in srv/chat-service.cds alongside getRelevantSteps / checkCode. Calls pathBetween + conceptsForUser; respects user's existing concept coverage when ranking; returns a numbered ordered list of tutorial slugs + titles + estimated time.
  4. Tool description for the LLM — well-tuned so the model picks this tool when the user asks for sequencing/ordering. Must NOT collide with the existing getRelevantSteps tool (which answers "what's the next step in this tutorial").
  5. Hybrid + smoke tests — test/hybrid/kg-path-between.test.js against a seeded fixture graph; test/smoke/joule-find-learning-path.test.js against the deployed chat endpoint.
  6. Telemetry — kg.joule.path_requested, kg.joule.path_returned events with pathLength, fromSlug, toSlug, userIdHash. Same dispatcher pattern as the sidebar's kg.sidebar.shown.

Out of scope for Phase 2

  • New UI surface (Joule chat is the surface; no /explore/ route)
  • "Why this path?" explanation rendering (would need a SPARQL trace explainer; defer to Phase 4 polish)
  • Multi-user collaborative paths
  • Path caching beyond the existing per-graph-version LRU pattern from Phase 1

Design questions to resolve

  • pathBetween semantics when no path exists — return empty array? Or compute a best-effort partial path through :relatedTo edges? Phase 2 design decision.
  • Path length cap — SPARQL property paths can be expensive; cap at 5 hops or 10? Probably 5 to keep query fast.
  • Multiple paths — return one (shortest), top-N (k-shortest), or all? Phase 2 picks one per the spec; Joule presents it.
  • Tool framing in the chat prompt — "find a path between A and B" vs "I want to learn X" require different parameter extraction. Probably two tool variants OR one with optional fromSlug (defaults to user's current position).

Risks

Risk Likelihood Mitigation
HANA KGE property-path support gaps Medium Day-1 spike (~1 day) on a test graph; fall back to JS-side traversal if needed
Joule tool overlap (collides with getRelevantSteps semantics) Medium Distinct tool name + descriptions tuned by adversarial test (LLM-judge picks correct tool ≥90% on a fixture)
pathBetween returns silly paths (e.g. via tagged-with link rather than meaningful prerequisites) High at first Restrict the SPARQL to property paths over :teaches/:requires only, NOT the full ontology. Tighten with hybrid-test corpus before shipping
Tool latency >2s makes chat UX laggy Low (SPARQL on this scale) Per-graph-version cache; LIMIT clauses; timeout from kg-sparql-client

Acceptance criteria

  • KnowledgeGraphService.pathBetween('cap-getting-started', 'cap-fiori-deployment') returns a non-empty ordered array of tutorial slugs
  • KnowledgeGraphService.conceptsForUser(userId) returns reasonable learned / partial for at least three real DEV users
  • Joule chat with prompt "I want to build a CAP service with Fiori UI" returns an ordered learning path
  • Telemetry events fire and land in UIEvent table
  • Hybrid + smoke tests pass on DEV
  • Phase 2 rollout note in docs/superpowers/done/ follows the Phase 1 template

Refs #381

Activity

  1. jung-thomas commented on Jun 23, 2026

    @jung-thomas
    ContributorAuthor

    PR #563 opens the implementation. Walkthrough of the design pivot for future readers:

    The issue body's canonical SPARQL ?a kg:teaches/^kg:requires*/kg:teaches ?b returned 0 rows on the v3 graph when probed pre-implementation — kg:requires is concept-level (not tutorial-level) and only 919 sparse edges exist (442 of 1357 concepts have any prereq). PR #563 ships a hybrid 3-arm pathBetween that uses what data IS dense:

    1. PREREQ (rank 1) — ?a kg:teaches/(^kg:requires)+/kg:teaches ?b — preferred when prereq edges exist
    2. CO_COMPLETED (rank 2) — ?a (kg:coCompletedWith)+ ?b — behavioral signal (~13k edges)
    3. SHARED_CONCEPT (rank 3) — ?a kg:teaches ?c. ?b kg:teaches ?c. — always-on

    Sorted by pathTypeRank ASC, LIMIT 10. JS layer dedups by slug, promotes the LLM-named toSlug when present, filters user-covered candidates (toSlug never dropped), hydrates with title + estimated time, renders a numbered markdown list.

    Future-friendly: when Phase 2.5 prereq-graph enrichment lands, ARM 1 will produce results more reliably without code changes here — the dispatcher already favors PREREQ.

    Acceptance criteria status

    • KnowledgeGraphService.pathBetween(fromSlug, toSlug) — wired via CDS action; gated by ChatSettings.kgPathBetweenEnabled
    • KnowledgeGraphService.conceptsForUser(userId) — wired via CDS action; helper at srv/lib/kg/concepts-for-user.js
    • Joule chat with "I want to build a CAP service with Fiori UI" returns an ordered learning path — requires admin to flip kgPathBetweenEnabled = true post-deploy
    • Telemetry events fire (kg.joule.path_requested, kg.joule.path_returned)
    • Hybrid + smoke tests pass — 14/14 hybrid, smoke skips cleanly when env vars unset
    • Phase 2 rollout note at docs/superpowers/done/ — file as follow-up; gets filled in 48h post-flag-flip

    Followups noted in the PR

    • Phase 2.5 prereq enrichment — separate issue to file; densifies the prereq sub-graph so ARM 1 produces results more reliably
    • kg:completedBy edge — Phase 4 architectural change for graph-side user identity; userIds currently stay out of the graph for privacy
    • Phase 3 Explore UI — Phase 3 — Explore page + concept landing pages #446 territory; pathBetween backend ready

    Latent-bug pair caught during execution

    Two pre-existing bugs from PR #555 surfaced when Phase 2 became the first production consumer of .response from the procedure layer:

    1. TaskRecords schema mismatch — Task 4's helper used TUTORIAL_ID but CAP uses (taskLegacyId, taskType). Caught by Task 8's hybrid test. Fixed in 885dc700.
    2. coerceRow doesn't unwrap { changes: [...] } — @cap-js/hana returns DO-block results in a wrapper, and the existing helper only unwrapped flat arrays. Every prior kgQuery()/kgAdminRunSparql() caller silently got { response: '' }. Caught when getConceptsForUser actually consumed .response. Fixed in b119f7df.

    Both fixes ship in this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions