Repository navigation
Knowledge graph of related tutorials and to help with learning discovery #381
Description
Activity
Brainstorm summary — knowledge graph of tutorials
Captured from a brainstorming session on 2026-06-17. The full design spec will follow in
docs/superpowers/specs/; this comment preserves the high-level direction and explicit Phase 2/3 scope.Direction (confirmed)
- Primary value: end-user discovery (mostly), with side benefit of better backend signal for existing surfaces, and serves as a showcase of SAP HANA Cloud's Knowledge Graph Engine (SPARQL/RDF triple store).
- Authority for the concept layer: AI-extracted from tutorial content. Admins curate (veto / merge / rename), they do not author concepts manually.
- HANA KGE is a projection, not source-of-truth. Canonical state lives in CDS-managed tables; the triple store is rebuilt from these on every concept change. This is the most important architectural decision — schema migrations, audit logging, draft semantics, and rebuild-from-scratch all "just work."
Ontology
8 predicates, mix of AI-extracted and structural-from-CDS:
Predicate Source Domain → Range :teachesAI Tutorial → Concept :requiresAI (high-confidence only) Concept → Concept :relatedToAI (co-occurrence + embedding) Concept → Concept :extendsAI ("if you've completed X" prose) Tutorial → Tutorial :partOfCDS Tutorial → Mission, Mission → Group :taggedWithCDS Tutorial → Tag :inCategoryCDS Mission → Category :aboutProductCDS ( software-product>*tags)Tutorial → Product :coCompletedWithanalytics (top-N, weighted) Tutorial → Tutorial Pipeline
Two cron jobs in
srv/jobs/:extractConcepts— nightly, per-tutorial, content-hash-keyed (mirrors AI-quiz cache pattern from AI-authored quizzes: generate [VALIDATE_N] candidates from tutorial body + rules.vr (follow-up to #171) #208). Constrained extraction: prompt includes existing concept registry; LLM proposes new concepts only for genuine gaps. Hard capKG_EXTRACT_BUILD_CAP=200per run. Steady-state cost ~$1-2/week.consolidateConcepts— weekly. Pairwise embedding similarity > 0.92 → auto-merge (preserve loser asstatus=MERGED,mergedIntopoints to canonical). DFS cycle detection on:requires→ auto-VETO weakest edge. Last step:graphRebuild— CLEAR + reload the named graph in HANA KGE from current CDS state.
Data model
Three new CDS entities in
db/knowledge-graph.cds:Concepts— canonical registry. Slug-keyed (@assert.unique), name + description, embedding centroid,status(ACTIVE/MERGED/VETOED),mergedIntoassociation for merge-don't-delete.TutorialConceptLinks— per-tutorial extracted concepts with confidence +contentHash(cache key) +modelVersion. One entity coversteachesandextendspredicates.ConceptEdges—:requiresand:relatedTobetween concepts, with confidence + LLM-cited evidence.
Query layer
New CAP service
KnowledgeGraphServiceat/graph(XSUAA-protected):- Named queries (typed, server-validated):
neighborhood(slug),pathBetween(from, to),conceptsForUser(userId). Same security model asAnalyticsService.runSelectQuery. runSparql(query)— admin-only, scopeKnowledgeGraph.Admin.- HANA KGE access via
EXECUTE STATEMENT 'SPARQL …'over the existingcds.connect.to('db')connection (no second client). - Per-slug result cache, invalidated by
graphVersionon rebuild.
The flagship Phase 1 SPARQL query is the 4-way UNION inside
neighborhood(slug)— multi-hop (teaches → requires → teaches), the kind of query that's awkward in SQL but elegant in SPARQL. This is the showcase moment.Phase 1 surfaces
Surface 1 — Tutorial sidebar (end-user-facing)
- New Vue 3 island at
hugo-apps/src/related-graph/ - Mounts on tutorial Object Page
- Four sections: This tutorial teaches, Prerequisites you might want first, Tutorials covering related concepts, What to learn next
- Behind feature flag
KNOWLEDGE_GRAPH_ENABLED(default OFF) - Hide-on-empty (no panel for tutorials with no extracted concepts yet)
- Clicks navigate to other tutorials; concept clicks are no-op in Phase 1 (Phase 3 routes to concept pages)
- Ranking for what to learn next combines SPARQL hop with
coCompletedWithweight + user concept coverage
Surface 2 — Admin concept review (
/admin-ui/#concepts-display)- Fiori Elements list page over
Concepts(peer of the existing 14 admin apps) - Capabilities: filter/search; inline edit name + description; Veto concept; Merge into… (value-help dialog over ACTIVE concepts); page-level Trigger graph rebuild
@cap-js/change-trackingenabled for audit trail- New scope
KnowledgeGraph.Adminadded to existingTutorial.Adminrole collection
Phase 1 ship-list
- 3 CDS entities (
Concepts,TutorialConceptLinks,ConceptEdges) - 1 CAP service (
KnowledgeGraphServicewith Phase 2 method stubs) - 2 lib files (
srv/lib/kg-extract.js,srv/lib/kg-queries.js) - 2 cron jobs (
extractConcepts,consolidateConcepts) - 1 Vue island (
hugo-apps/src/related-graph/) - 1 admin app (
app/admin/concepts/) + admin-shell side-nav entry - 1 feature flag (
KNOWLEDGE_GRAPH_ENABLED) - 1 XSUAA scope (
KnowledgeGraph.Admin) - 1-day spike: validate HANA KGE access via
EXECUTE STATEMENT 'SPARQL …'before locking the implementation
Risks
Risk Likelihood Mitigation AI extracts garbage concepts High in early runs Constrained extraction; admin veto; concept review tool ships in Phase 1 HANA KGE under-documented; auth/syntax surprises Medium Day-1 spike; fall back to REST endpoint if EXECUTE STATEMENTfailsPhase 1 wow factor lower than expected Medium Admin Concepts review is genuinely interesting on its own; SPARQL endpoint is the technical-credibility piece LLM costs balloon Low Hard build cap, content-hash cache, separate consolidation budget Cycles in :requiresexplode SPARQL property-path queriesMedium DFS validation, auto-VETO weakest edge HDI deploy wipes ConceptsdataMedium Snapshot row counts before deploy; registry is rebuildable from cache srv-qacp-list misses new lib filesMedium (recurring) PR-time audit + QA boot smoke Sidebar bloats tutorial OP load Low Lazy-load below fold, ETag cache, hide-on-empty Future scope (explicitly OUT of Phase 1)
- Phase 2 — Joule learning-path generator: new chat tool, NL → SPARQL → ordered tutorial path. Uses
pathBetween()andconceptsForUser()named queries already declared in the service shape. No new UI route — lives in existing Joule chat surface. Strong demo: "ask Joule for a learning path → see SPARQL → see KGE answer." - Phase 3 — Explore page: new
/explore/route. Force-directed (or constellation-style) graph viz. Tutorials, concepts, products as nodes; 8 predicates as typed edges. Click a node, graph re-centers. "Find a path from where I am to where I want to be" feature. Highest visual impact for the showcase but biggest scope. - Concept landing pages (
/concepts/<slug>/) — Phase 3 - Manual concept creation in admin — never; the showcase narrative is "AI builds the graph"
- Cross-corpus federation (e.g., SAP Help portal RDF alongside the tutorials graph) — interesting future
- Multi-language concept extraction — corpus is English-only today
- Real-time graph updates — graph is rebuilt per cron, not on every tutorial publish
- Embedding clustering as an alternative extraction strategy — considered (option B in design questions); rejected in favour of constrained per-tutorial extraction (option C). Could be revisited if vocabulary drift is bad
Decisions made (with rationale)
- AI-extracted (option C from question 3) over hand-curated ontology — showcases two SAP technologies (KGE + AI Core), matches author-self-service preference, and the consolidation job becomes its own demoable artifact.
- CAP cron job (option B from question 4) over CI-time extraction — independent of content publishing, content-hash cache survives across instances.
- Constrained per-tutorial extraction (option C from question 5) over corpus-wide clustering — caches well, produces coherent vocabulary, lets the registry stabilize naturally around ~80-150 concepts.
- All 8 predicates including the risky
:requires(mitigated by confidence threshold + cycle detection). - Phasing P1.1 — backend + sidebar (B) over Joule-first or admin-only. Smallest end-user surface, real eyeballs surface extraction-quality bugs we would never find from internal review.
- Named queries on the public surface, raw SPARQL admin-only. Same model as
AnalyticsService.runSelectQuery. whatToLearnNextranking happens in JS after the SPARQL hop, not in SPARQL itself. Keeps SPARQL clean; ranking stays where it is easy to tune.- HANA KGE access via
EXECUTE STATEMENT 'SPARQL …'over existingdb.run()connection, not a separate REST client. - Single
TutorialConceptLinksentity with predicate column over splitTeachesLinks+ExtendsLinksentities — same extraction pipeline, same cache key. @assert.uniqueonConcepts.slugfrom day one (PR fix: merge duplicate slugs + add @assert.unique guardrail #386 lesson learned).- Hide-on-empty sidebar rather than empty-state UI — Phase 1 ships only after first cron run.
KnowledgeGraph.Adminscope added to existingTutorial.Adminrole collection — no new role assignment work.
Next step: write the full design spec to
docs/superpowers/specs/2026-06-17-knowledge-graph-design.md, run it through the spec review loop, then implementation plan viasuperpowers:writing-plans.Nora had an excellent future suggestion:
Would be awesome to also add blog posts, learning journeys, basic trials and other relevant upskilling information
- added 11 commits that reference this issue
on Jun 18, 2026 14 remaining items
- added 15 commits that reference this issue
on Jun 19, 2026
Knowledge graph of related tutorials and to help with learning discovery