Skip to content

Knowledge graph of related tutorials and to help with learning discovery #381

Description

@jung-thomas

Knowledge graph of related tutorials and to help with learning discovery

Activity

  1. self-assigned this
    on Jun 17, 2026
  2. jung-thomas commented on Jun 17, 2026

    @jung-thomas
    ContributorAuthor

    Brainstorm summary — knowledge graph of tutorials

    Captured from a brainstorming session on 2026-06-17. The full design spec will follow in docs/superpowers/specs/; this comment preserves the high-level direction and explicit Phase 2/3 scope.

    Direction (confirmed)

    • Primary value: end-user discovery (mostly), with side benefit of better backend signal for existing surfaces, and serves as a showcase of SAP HANA Cloud's Knowledge Graph Engine (SPARQL/RDF triple store).
    • Authority for the concept layer: AI-extracted from tutorial content. Admins curate (veto / merge / rename), they do not author concepts manually.
    • HANA KGE is a projection, not source-of-truth. Canonical state lives in CDS-managed tables; the triple store is rebuilt from these on every concept change. This is the most important architectural decision — schema migrations, audit logging, draft semantics, and rebuild-from-scratch all "just work."

    Ontology

    8 predicates, mix of AI-extracted and structural-from-CDS:

    Predicate Source Domain → Range
    :teaches AI Tutorial → Concept
    :requires AI (high-confidence only) Concept → Concept
    :relatedTo AI (co-occurrence + embedding) Concept → Concept
    :extends AI ("if you've completed X" prose) Tutorial → Tutorial
    :partOf CDS Tutorial → Mission, Mission → Group
    :taggedWith CDS Tutorial → Tag
    :inCategory CDS Mission → Category
    :aboutProduct CDS (software-product>* tags) Tutorial → Product
    :coCompletedWith analytics (top-N, weighted) Tutorial → Tutorial

    Pipeline

    Two cron jobs in srv/jobs/:

    1. extractConcepts — nightly, per-tutorial, content-hash-keyed (mirrors AI-quiz cache pattern from AI-authored quizzes: generate [VALIDATE_N] candidates from tutorial body + rules.vr (follow-up to #171) #208). Constrained extraction: prompt includes existing concept registry; LLM proposes new concepts only for genuine gaps. Hard cap KG_EXTRACT_BUILD_CAP=200 per run. Steady-state cost ~$1-2/week.
    2. consolidateConcepts — weekly. Pairwise embedding similarity > 0.92 → auto-merge (preserve loser as status=MERGED, mergedInto points to canonical). DFS cycle detection on :requires → auto-VETO weakest edge. Last step: graphRebuild — CLEAR + reload the named graph in HANA KGE from current CDS state.

    Data model

    Three new CDS entities in db/knowledge-graph.cds:

    • Concepts — canonical registry. Slug-keyed (@assert.unique), name + description, embedding centroid, status (ACTIVE/MERGED/VETOED), mergedInto association for merge-don't-delete.
    • TutorialConceptLinks — per-tutorial extracted concepts with confidence + contentHash (cache key) + modelVersion. One entity covers teaches and extends predicates.
    • ConceptEdges — :requires and :relatedTo between concepts, with confidence + LLM-cited evidence.

    Query layer

    New CAP service KnowledgeGraphService at /graph (XSUAA-protected):

    • Named queries (typed, server-validated): neighborhood(slug), pathBetween(from, to), conceptsForUser(userId). Same security model as AnalyticsService.runSelectQuery.
    • runSparql(query) — admin-only, scope KnowledgeGraph.Admin.
    • HANA KGE access via EXECUTE STATEMENT 'SPARQL …' over the existing cds.connect.to('db') connection (no second client).
    • Per-slug result cache, invalidated by graphVersion on rebuild.

    The flagship Phase 1 SPARQL query is the 4-way UNION inside neighborhood(slug) — multi-hop (teaches → requires → teaches), the kind of query that's awkward in SQL but elegant in SPARQL. This is the showcase moment.

    Phase 1 surfaces

    Surface 1 — Tutorial sidebar (end-user-facing)

    • New Vue 3 island at hugo-apps/src/related-graph/
    • Mounts on tutorial Object Page
    • Four sections: This tutorial teaches, Prerequisites you might want first, Tutorials covering related concepts, What to learn next
    • Behind feature flag KNOWLEDGE_GRAPH_ENABLED (default OFF)
    • Hide-on-empty (no panel for tutorials with no extracted concepts yet)
    • Clicks navigate to other tutorials; concept clicks are no-op in Phase 1 (Phase 3 routes to concept pages)
    • Ranking for what to learn next combines SPARQL hop with coCompletedWith weight + user concept coverage

    Surface 2 — Admin concept review (/admin-ui/#concepts-display)

    • Fiori Elements list page over Concepts (peer of the existing 14 admin apps)
    • Capabilities: filter/search; inline edit name + description; Veto concept; Merge into… (value-help dialog over ACTIVE concepts); page-level Trigger graph rebuild
    • @cap-js/change-tracking enabled for audit trail
    • New scope KnowledgeGraph.Admin added to existing Tutorial.Admin role collection

    Phase 1 ship-list

    • 3 CDS entities (Concepts, TutorialConceptLinks, ConceptEdges)
    • 1 CAP service (KnowledgeGraphService with Phase 2 method stubs)
    • 2 lib files (srv/lib/kg-extract.js, srv/lib/kg-queries.js)
    • 2 cron jobs (extractConcepts, consolidateConcepts)
    • 1 Vue island (hugo-apps/src/related-graph/)
    • 1 admin app (app/admin/concepts/) + admin-shell side-nav entry
    • 1 feature flag (KNOWLEDGE_GRAPH_ENABLED)
    • 1 XSUAA scope (KnowledgeGraph.Admin)
    • 1-day spike: validate HANA KGE access via EXECUTE STATEMENT 'SPARQL …' before locking the implementation

    Risks

    Risk Likelihood Mitigation
    AI extracts garbage concepts High in early runs Constrained extraction; admin veto; concept review tool ships in Phase 1
    HANA KGE under-documented; auth/syntax surprises Medium Day-1 spike; fall back to REST endpoint if EXECUTE STATEMENT fails
    Phase 1 wow factor lower than expected Medium Admin Concepts review is genuinely interesting on its own; SPARQL endpoint is the technical-credibility piece
    LLM costs balloon Low Hard build cap, content-hash cache, separate consolidation budget
    Cycles in :requires explode SPARQL property-path queries Medium DFS validation, auto-VETO weakest edge
    HDI deploy wipes Concepts data Medium Snapshot row counts before deploy; registry is rebuildable from cache
    srv-qa cp-list misses new lib files Medium (recurring) PR-time audit + QA boot smoke
    Sidebar bloats tutorial OP load Low Lazy-load below fold, ETag cache, hide-on-empty

    Future scope (explicitly OUT of Phase 1)

    • Phase 2 — Joule learning-path generator: new chat tool, NL → SPARQL → ordered tutorial path. Uses pathBetween() and conceptsForUser() named queries already declared in the service shape. No new UI route — lives in existing Joule chat surface. Strong demo: "ask Joule for a learning path → see SPARQL → see KGE answer."
    • Phase 3 — Explore page: new /explore/ route. Force-directed (or constellation-style) graph viz. Tutorials, concepts, products as nodes; 8 predicates as typed edges. Click a node, graph re-centers. "Find a path from where I am to where I want to be" feature. Highest visual impact for the showcase but biggest scope.
    • Concept landing pages (/concepts/<slug>/) — Phase 3
    • Manual concept creation in admin — never; the showcase narrative is "AI builds the graph"
    • Cross-corpus federation (e.g., SAP Help portal RDF alongside the tutorials graph) — interesting future
    • Multi-language concept extraction — corpus is English-only today
    • Real-time graph updates — graph is rebuilt per cron, not on every tutorial publish
    • Embedding clustering as an alternative extraction strategy — considered (option B in design questions); rejected in favour of constrained per-tutorial extraction (option C). Could be revisited if vocabulary drift is bad

    Decisions made (with rationale)

    1. AI-extracted (option C from question 3) over hand-curated ontology — showcases two SAP technologies (KGE + AI Core), matches author-self-service preference, and the consolidation job becomes its own demoable artifact.
    2. CAP cron job (option B from question 4) over CI-time extraction — independent of content publishing, content-hash cache survives across instances.
    3. Constrained per-tutorial extraction (option C from question 5) over corpus-wide clustering — caches well, produces coherent vocabulary, lets the registry stabilize naturally around ~80-150 concepts.
    4. All 8 predicates including the risky :requires (mitigated by confidence threshold + cycle detection).
    5. Phasing P1.1 — backend + sidebar (B) over Joule-first or admin-only. Smallest end-user surface, real eyeballs surface extraction-quality bugs we would never find from internal review.
    6. Named queries on the public surface, raw SPARQL admin-only. Same model as AnalyticsService.runSelectQuery.
    7. whatToLearnNext ranking happens in JS after the SPARQL hop, not in SPARQL itself. Keeps SPARQL clean; ranking stays where it is easy to tune.
    8. HANA KGE access via EXECUTE STATEMENT 'SPARQL …' over existing db.run() connection, not a separate REST client.
    9. Single TutorialConceptLinks entity with predicate column over split TeachesLinks + ExtendsLinks entities — same extraction pipeline, same cache key.
    10. @assert.unique on Concepts.slug from day one (PR fix: merge duplicate slugs + add @assert.unique guardrail #386 lesson learned).
    11. Hide-on-empty sidebar rather than empty-state UI — Phase 1 ships only after first cron run.
    12. KnowledgeGraph.Admin scope added to existing Tutorial.Admin role collection — no new role assignment work.

    Next step: write the full design spec to docs/superpowers/specs/2026-06-17-knowledge-graph-design.md, run it through the spec review loop, then implementation plan via superpowers:writing-plans.

  3. jung-thomas commented on Jun 18, 2026

    @jung-thomas
    ContributorAuthor

    Nora had an excellent future suggestion:

    Would be awesome to also add blog posts, learning journeys, basic trials and other relevant upskilling information

  4. 14 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions