Feature request: --backend claude-cli — route through the local Claude Code CLI instead of requiring ANTHROPIC_API_KEY
Package: graphifyy v0.7.16
Use case: Claude Code Pro/Max subscribers running graphify's semantic pass on their own codebase.
What
Add a claude-cli backend that shells out to the locally-installed claude CLI (Claude Code) via claude -p --output-format json instead of calling the Anthropic API directly.
Why
--backend claude currently requires ANTHROPIC_API_KEY — pay-as-you-go API billing, separate from the Claude Code subscription. For Pro/Max subscribers, this means:
- Provisioning a second Anthropic account (or a billing key on the same one) just to run graphify's one-off semantic pass.
- Paying twice — once via the subscription (which they're already using for Claude Code itself) and once via API credit (for graphify).
The local claude CLI is already authenticated via the user's subscription OAuth and supports a non-interactive -p mode with --output-format json that returns a parseable envelope including the model result, usage stats, stop reason, and modelUsage map. It can serve as a drop-in transport for graphify's _call_claude() use case.
Verified on a real codebase: ran graphify's full semantic extract (167k input / 8k output tokens, Opus 4.7 1M context) entirely through claude -p, billed against my Max plan instead of API credit.
Approach
Mirror the existing claude backend with a new claude-cli arm:
- New
BACKENDS["claude-cli"] entry with env_key: None and zero pricing (calls bill against the plan, not the API — graphify's cost estimator treats this as $0).
- New
_call_claude_cli(user_message, max_tokens) that:
- Verifies
shutil.which("claude") finds the binary.
- Subprocess-calls
claude -p --output-format json --append-system-prompt <_EXTRACTION_SYSTEM> with the user message piped on stdin.
- Parses the JSON envelope: extracts
result (the model output), sums usage.input_tokens + cache_read_input_tokens + cache_creation_input_tokens for the input token total, and maps stop_reason == "max_tokens" to graphify's finish_reason == "length".
- New dispatch arm in
extract_files_direct plus the lower _call_llm helper.
- New env-key carve-out in
__main__.py extract validation: when backend == "claude-cli", check shutil.which("claude") instead of an env variable. Mirrors ollama's local-loopback carve-out.
The user must have authenticated claude interactively at least once so the OAuth tokens are stored locally; after that, -p mode runs against the subscription without further prompts.
Notes for the maintainer
- The current Claude Code CLI's
--bare mode (which would strip the default system prompt) requires ANTHROPIC_API_KEY and would defeat the purpose, so we use --append-system-prompt instead. The base Claude Code prompt is appended, but the _EXTRACTION_SYSTEM prompt is strict enough that the model produces valid JSON anyway. Verified empirically — see PR for test fixture using a real Claude Code response envelope.
- Tests in the PR mock
subprocess.run and shutil.which so the suite still runs on CI without the claude binary or network access.
- Performance is comparable to the API backend: the Claude Code CLI caches its system prompt server-side, so subsequent calls in a session benefit from the same cache discount as direct API calls.
PR with implementation + tests incoming.
Feature request:
--backend claude-cli— route through the local Claude Code CLI instead of requiringANTHROPIC_API_KEYPackage:
graphifyyv0.7.16Use case: Claude Code Pro/Max subscribers running graphify's semantic pass on their own codebase.
What
Add a
claude-clibackend that shells out to the locally-installedclaudeCLI (Claude Code) viaclaude -p --output-format jsoninstead of calling the Anthropic API directly.Why
--backend claudecurrently requiresANTHROPIC_API_KEY— pay-as-you-go API billing, separate from the Claude Code subscription. For Pro/Max subscribers, this means:The local
claudeCLI is already authenticated via the user's subscription OAuth and supports a non-interactive-pmode with--output-format jsonthat returns a parseable envelope including the model result, usage stats, stop reason, and modelUsage map. It can serve as a drop-in transport for graphify's_call_claude()use case.Verified on a real codebase: ran graphify's full semantic extract (167k input / 8k output tokens, Opus 4.7 1M context) entirely through
claude -p, billed against my Max plan instead of API credit.Approach
Mirror the existing
claudebackend with a newclaude-cliarm:BACKENDS["claude-cli"]entry withenv_key: Noneand zero pricing (calls bill against the plan, not the API — graphify's cost estimator treats this as $0)._call_claude_cli(user_message, max_tokens)that:shutil.which("claude")finds the binary.claude -p --output-format json --append-system-prompt <_EXTRACTION_SYSTEM>with the user message piped on stdin.result(the model output), sumsusage.input_tokens + cache_read_input_tokens + cache_creation_input_tokensfor the input token total, and mapsstop_reason == "max_tokens"to graphify'sfinish_reason == "length".extract_files_directplus the lower_call_llmhelper.__main__.pyextract validation: whenbackend == "claude-cli", checkshutil.which("claude")instead of an env variable. Mirrors ollama's local-loopback carve-out.The user must have authenticated
claudeinteractively at least once so the OAuth tokens are stored locally; after that,-pmode runs against the subscription without further prompts.Notes for the maintainer
--baremode (which would strip the default system prompt) requiresANTHROPIC_API_KEYand would defeat the purpose, so we use--append-system-promptinstead. The base Claude Code prompt is appended, but the_EXTRACTION_SYSTEMprompt is strict enough that the model produces valid JSON anyway. Verified empirically — see PR for test fixture using a real Claude Code response envelope.subprocess.runandshutil.whichso the suite still runs on CI without theclaudebinary or network access.PR with implementation + tests incoming.