Compare commits

..
Author SHA1 Message Date
william 6d9d081030 Merge pull request 'Remove claude-agent Matrix presence — Hermes only' (#14) from remove/claude-bot-matrix-presence into main
build-agent / build-and-push (push) Successful in 10s
Reviewed-on: #14
2026-08-23 16:16:58 +00:00
william 6bc862e051 Remove claude-agent's Matrix presence entirely — Hermes is the only agent in Matrix
By request: one agent in Matrix, not several. Removes matrixBot.js,
router.js (chat-vs-code-task classifier), litellm.js (claude-agent's own
LiteLLM client), the matrix-bot-sdk dependency, runChatTask() and its
gitea.js branch/PR helpers (createBranch/createPullRequest — only ever
called from the now-removed chat flow), and every Matrix/LiteLLM env var
from claude-agent's compose service.

claude-agent already left the control room manually before this merge.
It keeps its Gitea-webhook-triggered PR review, which never touched Matrix
or LiteLLM to begin with.

Makes PR #12 (the claude-bot/Hermes cross-reply cascade fix) moot — the
bug can't happen once claude-agent has no Matrix client at all. Close #12
without merging once this lands.
2026-08-23 16:16:19 +00:00
william fff021283d Merge pull request 'Re-add claude-subscription LiteLLM route — confirmed working for the real CLI' (#13) from feat/litellm-claude-subscription-route into main
Reviewed-on: #13
2026-08-23 16:09:47 +00:00
william f4f785e0bb Re-add the claude-subscription LiteLLM route — confirmed working for the real CLI
Earlier this session I removed this route after a curl-based test got
rejected by Anthropic and concluded OAuth subscription forwarding doesn't
work through a proxy at all. That conclusion was wrong: the real `claude`
CLI binary, with ANTHROPIC_BASE_URL pointed at litellm, successfully
completed a request billed against the subscription. The earlier curl test
just didn't replicate whatever header/fingerprint Anthropic requires from
genuine Claude Code CLI traffic — LiteLLM relays that fine when the real
CLI is the caller, but a hand-built request from any other client (Hermes
included) still gets rejected the same way curl did.
2026-08-23 16:08:26 +00:00
william c88fdcc2ea Merge pull request 'Fix Hermes: gateway command + disable network-reachable API server' (#11) from fix/hermes-gateway-command-and-api-server into main
Reviewed-on: #11
2026-08-23 15:52:35 +00:00
william 98762e764a Fix Hermes container: add gateway command, disable network-reachable API server
Without an explicit command the image launches the interactive CLI by
default, which immediately exits with 'Input is not a terminal' in a
detached container — it was doing nothing on every restart. Also disables
API_SERVER_ENABLED: Hermes itself warns at startup that a network-reachable
API server combined with the default unsandboxed 'local' terminal backend
gives any caller on the network full terminal/file access. Not needed yet
(Matrix is the actual interface) — can re-enable properly (with a sandboxed
terminal backend) if claude-agent ever needs to call Hermes programmatically.
2026-08-23 15:52:08 +00:00
william 5507192ee6 Merge pull request 'Add Hermes Agent as a second native Matrix presence' (#10) from feat/hermes-matrix into main
build-agent / build-and-push (push) Successful in 11s
Reviewed-on: #10
2026-08-23 15:49:05 +00:00
william 9fa025f7d5 Add Hermes Agent as a second native Matrix presence
Deployed as its own service (pinned nousresearch/hermes-agent:v2026.8.19),
own Matrix bot account (@hermes), own OpenRouter-backed model config, and
its own OpenAI-compatible API server (internal network only, for possible
future use by claude-agent). Joins the same control room but only responds
when explicitly @mentioned, restricted to the human user — no conflict with
claude-bot's default no-prefix chat routing. claude-agent's router now
ignores messages addressed to @hermes so both bots don't answer the same
message.

Bridge networking (the 'web' network), not the image's default host mode —
no reason for an agent container to share the host's network namespace when
everything it needs (the homeserver, OpenRouter) is reachable over the
existing bridge.
2026-08-23 15:46:58 +00:00
12 changed files with 94 additions and 290 deletions
+12 -15
View File
@@ -21,25 +21,22 @@ GITEA_REGISTRY_IMAGE=gitea.apps.williamturner.eu/<your-gitea-username>/<repo-nam
# to generate this — it's a long-lived OAuth token, not an API key. # to generate this — it's a long-lived OAuth token, not an API key.
CLAUDE_CODE_OAUTH_TOKEN= CLAUDE_CODE_OAUTH_TOKEN=
# --- matrix bot --- # --- litellm (local LLM gateway — used by Hermes, see litellm-config.yaml) ---
MATRIX_HOMESERVER_URL=https://matrix.apps.williamturner.eu
MATRIX_BOT_TOKEN=
MATRIX_CONTROL_ROOM_ID=
# The bot's own Matrix ID (@username:server), e.g. @claude-bot:matrix.apps.williamturner.eu
# — set explicitly rather than fetched via the API (that call 404s against Continuwuity).
# Required: without it the bot can't tell its own messages apart from real ones and would
# reply to itself in a loop, so it refuses to start.
MATRIX_BOT_USER_ID=
# Comma-separated "owner/repo" list the chat router is allowed to open code-change PRs
# against. A plain chat message mentioning a repo NOT in this list is treated as chat,
# never as a code task — the router only matches confidently against known repos.
KNOWN_REPOS=william/gitops-automation
# --- litellm (local LLM gateway — see litellm-config.yaml) ---
OPENROUTER_API_KEY= OPENROUTER_API_KEY=
# Any random string; also used as litellm's general_settings.master_key. # Any random string; also used as litellm's general_settings.master_key.
LITELLM_MASTER_KEY= LITELLM_MASTER_KEY=
# --- hermes (the only agent with a Matrix presence — see README) ---
# Your own Matrix ID — Hermes only responds to this user, and only when @mentioned
# in a room (free-response in DMs).
MATRIX_HUMAN_USER_ID=@william:matrix.apps.williamturner.eu
# Access token for the @hermes bot account — register it on the homeserver, then log
# in as it via /_matrix/client/v3/login to get this token (see README).
HERMES_MATRIX_ACCESS_TOKEN=
# Any random string — bearer key for Hermes's own OpenAI-compatible API server
# (internal network only, not published anywhere).
HERMES_API_SERVER_KEY=
# --- portainer (GitOps redeploy) --- # --- portainer (GitOps redeploy) ---
PORTAINER_STACK_WEBHOOK_URL= PORTAINER_STACK_WEBHOOK_URL=
-8
View File
@@ -1,8 +0,0 @@
# Contributing
## Chat routing
Matrix control room messages no longer need a `!claude` or `!ai` command
prefix. Every message is routed automatically (via the LiteLLM gateway and
`agent/src/router.js`'s classifier) to the right handler — no prefix is
required or recognized anymore.
+25 -21
View File
@@ -1,21 +1,25 @@
# gitops-automation # gitops-automation
Claude Code automation wired into Gitea + Portainer + Matrix on this VPS. See Claude Code + Hermes automation wired into Gitea + Portainer + Matrix on this VPS. See
`~/.claude/plans/cozy-honking-lantern.md` on the host for the full design rationale. `~/.claude/plans/cozy-honking-lantern.md` on the host for the original design rationale
(some of it — the Matrix chat bot on claude-agent — has since been superseded, see below).
Three things this gives you: What this gives you today:
- **PR review**: opening/updating a PR in a watched Gitea repo gets a Claude-authored - **PR review**: opening/updating a PR in a watched Gitea repo gets a Claude-authored
review comment. review comment (`claude-agent`, triggered by a Gitea webhook — nothing to do with Matrix).
- **GitOps redeploy**: pushing to `main` on this repo rebuilds the `claude-agent` image - **GitOps redeploy**: pushing to `main` on this repo rebuilds the `claude-agent` image
(Gitea Actions) and redeploys the stack (Portainer webhook). (Gitea Actions) and redeploys the stack, chained as the last step of that same workflow.
- **Chat-driven coding agent**: `!claude owner/repo <instruction>` in the Matrix control - **Matrix**: **Hermes is the only agent present in Matrix.** `claude-agent` used to also
room clones the repo, runs Claude Code, and opens a PR with the result. run a Matrix bot (`!claude`/`!ai` commands, then no-prefix auto-routing) — that's been
- **Ask other models**: `!ai <prompt>` (default model) or `!ai provider/model <prompt>` removed entirely (by request: one agent in Matrix, not several). Hermes has its own
(e.g. `!ai google/gemini-2.0-flash-001 explain this error`) queries any model on native Matrix connection, responds to `@hermes <message>` in shared rooms (no mention
OpenRouter and replies in the room. No repo/file access — just a chat reply, unlike needed in DMs), and has its own tools (terminal, code execution, web search, etc.) — see
`!claude` which is the only command that can edit files and open PRs. its docs at https://hermes-agent.nousresearch.com for what it can do. It does not (yet)
have the old branch/PR-opening workflow the Matrix bot used to have; that logic still
exists in git history if it's worth reviving as a Hermes tool/skill later.
Nothing here auto-merges. Every path stops at a comment or an open PR — a human clicks merge. Nothing here auto-merges PRs. Every path stops at a comment or an open PR — a human clicks
merge (enforced by branch protection on `main`, not just by convention — see below).
## Prerequisites (one-time, on the VPS) ## Prerequisites (one-time, on the VPS)
@@ -125,17 +129,17 @@ docker network create web
## Smoke test ## Smoke test
- Open a throwaway PR on a repo with the PR webhook set → expect a Claude review comment. - Open a throwaway PR on a repo with the PR webhook set → expect a Claude review comment.
- In the Matrix control room: `!claude owner/repo add a comment to the README` → expect a - In the Matrix control room: `@hermes hello` → expect a reply from Hermes.
"working on it" reply, then a PR link. - `git push` to `main` on this repo (touching `agent/**`) → expect a Gitea Actions run,
- `git push` to `main` on this repo → expect a Gitea Actions run, then a Portainer redeploy. then a chained Portainer redeploy at the end of that same workflow.
## Notes ## Notes
- `agent/src/runner.js` is the only thing that ever runs `git commit`/`git push`/`git - `agent/src/runner.js` only ever runs read-only `git clone`/`fetch`/`checkout` for the PR
checkout` — Claude Code itself is explicitly denied those tools (`--disallowedTools`), diff it reviews — Claude Code itself is denied `Edit`/`Write`/commit/push tools
so even a misbehaving prompt can't push directly or touch `main`. (`--disallowedTools`), so this path can never modify a repo, only comment on it.
- The Matrix bot only reacts inside `MATRIX_CONTROL_ROOM_ID`; keep that room invite-only.
- `.gitea/workflows/build.yml` assumes the act_runner label `docker` — check - `.gitea/workflows/build.yml` assumes the act_runner label `docker` — check
`GITEA_RUNNER_LABELS` in `docker-compose.yml` matches what you actually registered. `GITEA_RUNNER_LABELS` in `docker-compose.yml` matches what you actually registered.
- Hermes's own config/memory/skills live in `/home/william/hermes-data` on the host
<!-- gitops loop smoke test 2026-08-23T11:23:34Z --> (bind-mounted, not in this repo) — back that up separately if it accumulates anything
worth keeping.
+1 -2
View File
@@ -8,7 +8,6 @@
"start": "node src/server.js" "start": "node src/server.js"
}, },
"dependencies": { "dependencies": {
"express": "^4.19.2", "express": "^4.19.2"
"matrix-bot-sdk": "^0.7.1"
} }
} }
-21
View File
@@ -24,27 +24,6 @@ export async function postPRComment(owner, repo, index, body) {
await assertOk(res, "post PR comment"); await assertOk(res, "post PR comment");
} }
export async function createBranch(owner, repo, newBranch, oldBranch = "main") {
const url = `${GITEA_URL}/api/v1/repos/${owner}/${repo}/branches`;
const res = await fetch(url, {
method: "POST",
headers: authHeaders(),
body: JSON.stringify({ new_branch_name: newBranch, old_branch_name: oldBranch }),
});
await assertOk(res, "create branch");
}
export async function createPullRequest(owner, repo, { head, base = "main", title, body }) {
const url = `${GITEA_URL}/api/v1/repos/${owner}/${repo}/pulls`;
const res = await fetch(url, {
method: "POST",
headers: authHeaders(),
body: JSON.stringify({ head, base, title, body }),
});
await assertOk(res, "create PR");
return res.json();
}
// Injects the agent's token into a Gitea clone URL so git operations don't need SSH keys. // Injects the agent's token into a Gitea clone URL so git operations don't need SSH keys.
export function authenticatedCloneUrl(cloneUrl) { export function authenticatedCloneUrl(cloneUrl) {
const u = new URL(cloneUrl); const u = new URL(cloneUrl);
-37
View File
@@ -1,37 +0,0 @@
const LITELLM_BASE_URL = process.env.LITELLM_BASE_URL || "http://litellm:4000";
const LITELLM_MASTER_KEY = process.env.LITELLM_MASTER_KEY;
// OpenAI-compatible chat completion, for OpenRouter-backed models routed through the
// local LiteLLM gateway (e.g. "auto" — OpenRouter's own prompt-aware auto-router).
export async function chatCompletion(model, prompt) {
if (!LITELLM_MASTER_KEY) {
throw new Error("LITELLM_MASTER_KEY is not set");
}
const res = await fetch(`${LITELLM_BASE_URL}/v1/chat/completions`, {
method: "POST",
headers: {
Authorization: `Bearer ${LITELLM_MASTER_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model,
messages: [{ role: "user", content: prompt }],
// Some models default max_tokens to their full context window (e.g. 65536), which
// can exceed available credit balance before a single token is generated. This is a
// quick chat reply, not a long-form task — cap it.
max_tokens: 1024,
}),
});
if (!res.ok) {
throw new Error(`LiteLLM request failed: ${res.status} ${await res.text()}`);
}
const data = await res.json();
const content = data.choices?.[0]?.message?.content;
if (!content) {
throw new Error(`LiteLLM returned no content: ${JSON.stringify(data)}`);
}
return content;
}
-75
View File
@@ -1,75 +0,0 @@
import { MatrixClient, SimpleFsStorageProvider } from "matrix-bot-sdk";
import { runChatTask } from "./runner.js";
import { routeMessage, chatReply } from "./router.js";
const HOMESERVER_URL = process.env.MATRIX_HOMESERVER_URL;
const ACCESS_TOKEN = process.env.MATRIX_BOT_TOKEN;
const CONTROL_ROOM_ID = process.env.MATRIX_CONTROL_ROOM_ID;
const GITEA_URL = process.env.GITEA_URL;
// Set explicitly rather than fetched via client.getUserId() — that call hits /whoami,
// which (like /joined_rooms before it) 404s against Continuwuity for reasons unrelated
// to the endpoint itself. This is also the only reliable way to filter the bot's own
// messages now that there's no command prefix to naturally exclude them by.
const BOT_USER_ID = process.env.MATRIX_BOT_USER_ID;
const KNOWN_REPOS = (process.env.KNOWN_REPOS || "")
.split(",")
.map((r) => r.trim())
.filter(Boolean);
const MAX_REPLY_LENGTH = 4000;
function truncate(text) {
if (text.length <= MAX_REPLY_LENGTH) return text;
return `${text.slice(0, MAX_REPLY_LENGTH)}\n\n[truncated]`;
}
export async function startMatrixBot() {
if (!HOMESERVER_URL || !ACCESS_TOKEN || !CONTROL_ROOM_ID) {
console.warn("Matrix env vars not set — skipping bot startup");
return;
}
if (!BOT_USER_ID) {
console.warn("MATRIX_BOT_USER_ID not set — bot could reply to its own messages, skipping startup");
return;
}
const storage = new SimpleFsStorageProvider("/workspace/matrix-bot-storage.json");
const client = new MatrixClient(HOMESERVER_URL, ACCESS_TOKEN, storage);
client.on("room.invite", async (roomId) => {
try {
await client.joinRoom(roomId);
} catch (err) {
console.error("failed to join invited room", roomId, err.message);
}
});
client.on("room.message", async (roomId, event) => {
if (roomId !== CONTROL_ROOM_ID) return;
if (event.sender === BOT_USER_ID) return;
const body = event.content?.body;
if (!body) return;
try {
const decision = await routeMessage(body, KNOWN_REPOS);
if (decision.type === "code_task") {
const [owner, repo] = decision.repo.split("/");
await client.sendText(roomId, `Working on it: ${decision.repo}${decision.instruction}`);
const cloneUrl = `${GITEA_URL}/${owner}/${repo}.git`;
const pr = await runChatTask({ owner, repo, cloneUrl, instruction: decision.instruction });
await client.sendText(roomId, `Opened PR: ${pr.html_url}`);
return;
}
const reply = await chatReply(body);
await client.sendText(roomId, truncate(reply));
} catch (err) {
console.error("message handling failed", err);
await client.sendText(roomId, `Failed: ${err.message}`);
}
});
await client.start();
console.log("Matrix bot started, room:", CONTROL_ROOM_ID);
}
-45
View File
@@ -1,45 +0,0 @@
import { chatCompletion } from "./litellm.js";
const ROUTER_MODEL = "router-classifier";
const CHAT_MODEL = "auto";
function systemPrompt(knownRepos) {
return [
"You are a routing classifier for a chat bot. Given a user message, decide whether it is:",
'- "chat": a question, discussion, or anything that just needs a text reply.',
'- "code_task": a request to change a specific code repository (add/edit/fix something)',
" where the repository is clearly one of the known repositories below.",
"",
`Known repositories: ${knownRepos.join(", ") || "(none configured)"}`,
"",
"Reply with ONLY a JSON object, nothing else:",
'{"type":"chat"}',
'or',
'{"type":"code_task","repo":"owner/repo","instruction":"clear imperative instruction"}',
"",
"If it sounds like a code change but you can't confidently match it to one of the known",
'repositories, reply {"type":"chat"} instead of guessing.',
].join("\n");
}
function parseDecision(raw) {
try {
const cleaned = raw.trim().replace(/^```(?:json)?\n?/, "").replace(/```$/, "");
const parsed = JSON.parse(cleaned);
if (parsed.type === "code_task" && parsed.repo && parsed.instruction) {
return parsed;
}
} catch {
// fall through to chat — an unparseable classification is not a reason to edit a repo
}
return { type: "chat" };
}
export async function routeMessage(text, knownRepos) {
const raw = await chatCompletion(ROUTER_MODEL, `${systemPrompt(knownRepos)}\n\nMessage: ${text}`);
return parseDecision(raw);
}
export async function chatReply(text) {
return chatCompletion(CHAT_MODEL, text);
}
+8 -42
View File
@@ -2,8 +2,7 @@ import { execFile } from "node:child_process";
import { promisify } from "node:util"; import { promisify } from "node:util";
import { mkdtemp, rm, mkdir } from "node:fs/promises"; import { mkdtemp, rm, mkdir } from "node:fs/promises";
import path from "node:path"; import path from "node:path";
import crypto from "node:crypto"; import { authenticatedCloneUrl } from "./gitea.js";
import { authenticatedCloneUrl, createBranch, createPullRequest } from "./gitea.js";
const execFileAsync = promisify(execFile); const execFileAsync = promisify(execFile);
const WORKSPACE_ROOT = "/workspace"; const WORKSPACE_ROOT = "/workspace";
@@ -29,10 +28,12 @@ async function withWorkspace(fn) {
// Runs Claude Code headless. Unattended containers have no TTY to answer permission // Runs Claude Code headless. Unattended containers have no TTY to answer permission
// prompts, so this trusts the sandboxing of the throwaway clone dir instead: // prompts, so this trusts the sandboxing of the throwaway clone dir instead:
// bypassPermissions to avoid hanging, plus --disallowedTools as defense in depth so // bypassPermissions to avoid hanging, plus --disallowedTools as defense in depth so
// Claude can never push/commit/checkout itself — this script owns those steps. // Claude can never push/commit/checkout, or edit files — this is read-only review.
async function runClaude(cwd, prompt, { allowEdits }) { async function runClaude(cwd, prompt) {
const disallowed = ["Bash(git push:*)", "Bash(git commit:*)", "Bash(git checkout:*)"]; const disallowed = [
if (!allowEdits) disallowed.push("Edit", "Write", "NotebookEdit"); "Bash(git push:*)", "Bash(git commit:*)", "Bash(git checkout:*)",
"Edit", "Write", "NotebookEdit",
];
const args = [ const args = [
"-p", prompt, "-p", prompt,
@@ -59,41 +60,6 @@ export async function reviewPullRequest({ owner, repo, ref, cloneUrl, prTitle, p
`PR description:\n${prBody}`, `PR description:\n${prBody}`,
].join("\n"); ].join("\n");
return runClaude(dir, prompt, { allowEdits: false }); return runClaude(dir, prompt);
});
}
export async function runChatTask({ owner, repo, cloneUrl, instruction }) {
return withWorkspace(async (dir) => {
const authedUrl = authenticatedCloneUrl(cloneUrl);
await run("git", ["clone", "--quiet", authedUrl, dir]);
const branch = `claude/${crypto.randomBytes(4).toString("hex")}`;
await createBranch(owner, repo, branch);
await run("git", ["fetch", "--quiet", "origin", branch], { cwd: dir });
await run("git", ["checkout", "--quiet", branch], { cwd: dir });
const prompt = [
"Implement the change described below in this repository. Make the smallest",
"correct change that satisfies it. Do not run git commit, git push, or git",
"checkout yourself — just edit files; committing and pushing happens separately.",
`Instruction: ${instruction}`,
].join("\n");
await runClaude(dir, prompt, { allowEdits: true });
await run("git", ["add", "-A"], { cwd: dir });
const status = await run("git", ["status", "--porcelain"], { cwd: dir });
if (!status.trim()) {
throw new Error("Claude made no changes for this instruction");
}
await run("git", ["commit", "-m", `claude: ${instruction}`.slice(0, 200)], { cwd: dir });
await run("git", ["push", "--quiet", "origin", branch], { cwd: dir });
return createPullRequest(owner, repo, {
head: branch,
title: `claude: ${instruction}`.slice(0, 200),
body: `Requested via Matrix:\n\n> ${instruction}`,
});
}); });
} }
-5
View File
@@ -2,7 +2,6 @@ import express from "express";
import crypto from "node:crypto"; import crypto from "node:crypto";
import { postPRComment } from "./gitea.js"; import { postPRComment } from "./gitea.js";
import { reviewPullRequest } from "./runner.js"; import { reviewPullRequest } from "./runner.js";
import { startMatrixBot } from "./matrixBot.js";
const app = express(); const app = express();
app.use( app.use(
@@ -60,7 +59,3 @@ app.post("/webhooks/gitea", async (req, res) => {
app.listen(PORT, () => { app.listen(PORT, () => {
console.log(`claude-agent listening on :${PORT}`); console.log(`claude-agent listening on :${PORT}`);
}); });
startMatrixBot().catch((err) => {
console.error("matrix bot failed to start:", err);
});
+35 -12
View File
@@ -61,10 +61,43 @@ services:
# Internal only — no Traefik labels. No reason to expose an LLM gateway holding a # Internal only — no Traefik labels. No reason to expose an LLM gateway holding a
# master key and OAuth-forwarding config to the public internet. # master key and OAuth-forwarding config to the public internet.
hermes:
# Pinned to a specific dated release, not :latest — same rationale as litellm above.
image: nousresearch/hermes-agent:v2026.8.19
container_name: hermes
restart: unless-stopped
environment:
HERMES_UID: "1000"
HERMES_GID: "1000"
# Internal container address, not the public HTTPS one — same docker network as
# matrix-homeserver, no reason to round-trip through Traefik/TLS for this.
MATRIX_HOMESERVER: http://matrix-homeserver:8008
MATRIX_ACCESS_TOKEN: ${HERMES_MATRIX_ACCESS_TOKEN}
# Only you can trigger it; and only with an explicit @hermes mention in shared
# rooms (DMs to it would respond unprompted, per Hermes's own default behavior).
MATRIX_ALLOWED_USERS: ${MATRIX_HUMAN_USER_ID}
MATRIX_REQUIRE_MENTION: "true"
OPENROUTER_API_KEY: ${OPENROUTER_API_KEY}
# Left disabled: Hermes itself warns that a network-reachable API server combined
# with the default unsandboxed ('local') terminal backend gives any caller full
# terminal/file access within the container. Matrix is the actual interface in use;
# re-enable (API_SERVER_HOST: 0.0.0.0) only alongside terminal.backend: docker if
# claude-agent ever needs to call Hermes programmatically.
API_SERVER_ENABLED: "false"
volumes:
- /home/william/hermes-data:/opt/data
networks:
- web
# Without this the image's default command launches the interactive CLI, which
# immediately exits ("Input is not a terminal") since a detached container has no
# stdin — the container then just sits there having done nothing, every restart.
command: ["gateway", "run"]
claude-agent: claude-agent:
# Gitea PR-review only now — no Matrix presence (see hermes above; only one agent
# is meant to be in Matrix). Still triggered by Gitea's pull_request webhook and
# posts review comments there, entirely independent of Matrix/LiteLLM.
image: ${GITEA_REGISTRY_IMAGE} image: ${GITEA_REGISTRY_IMAGE}
depends_on:
- litellm
container_name: claude-agent container_name: claude-agent
restart: unless-stopped restart: unless-stopped
# Explicit vars, not env_file: .env — Portainer's git-based stack deploy clones the # Explicit vars, not env_file: .env — Portainer's git-based stack deploy clones the
@@ -77,16 +110,6 @@ services:
# Claude subscription (Pro/Max) auth via `claude setup-token`, not API billing — # Claude subscription (Pro/Max) auth via `claude setup-token`, not API billing —
# Claude Code reads this in preference to ANTHROPIC_API_KEY when both could apply. # Claude Code reads this in preference to ANTHROPIC_API_KEY when both could apply.
CLAUDE_CODE_OAUTH_TOKEN: ${CLAUDE_CODE_OAUTH_TOKEN} CLAUDE_CODE_OAUTH_TOKEN: ${CLAUDE_CODE_OAUTH_TOKEN}
MATRIX_HOMESERVER_URL: ${MATRIX_HOMESERVER_URL}
MATRIX_BOT_TOKEN: ${MATRIX_BOT_TOKEN}
MATRIX_CONTROL_ROOM_ID: ${MATRIX_CONTROL_ROOM_ID}
MATRIX_BOT_USER_ID: ${MATRIX_BOT_USER_ID}
KNOWN_REPOS: ${KNOWN_REPOS}
# All model calls now go through the local litellm service, not OpenRouter directly —
# one gateway for OpenRouter's models (incl. its auto-router) and, for the
# claude-subscription route, Anthropic itself via the forwarded OAuth token above.
LITELLM_BASE_URL: http://litellm:4000
LITELLM_MASTER_KEY: ${LITELLM_MASTER_KEY}
volumes: volumes:
- agent_workspace:/workspace - agent_workspace:/workspace
networks: networks:
+13 -7
View File
@@ -11,13 +11,19 @@ model_list:
model: openrouter/openai/gpt-4o-mini model: openrouter/openai/gpt-4o-mini
api_key: os.environ/OPENROUTER_API_KEY api_key: os.environ/OPENROUTER_API_KEY
# NOT included: a "claude-subscription" route forwarding the Claude Pro/Max OAuth token # Routes to Anthropic using the CALLER's forwarded Authorization header (the Claude
# (from `claude setup-token`) through to Anthropic's raw API. Tested and confirmed # Pro/Max subscription OAuth token) instead of a LiteLLM-held API key — billed against
# non-functional — Anthropic returns a generic rate_limit_error for ANY direct API call # the subscription, not per-token. CONFIRMED WORKING, but only for the real `claude`
# using this token type outside the real Claude Code CLI client (reproduced with plain # CLI binary as caller (tested: `claude -p` with ANTHROPIC_BASE_URL pointed here
# curl straight to api.anthropic.com, bypassing LiteLLM entirely, same result). The # returned a real completion). An earlier test with plain curl replicating the same
# subscription token only works through the actual Claude Code CLI, which is what # request shape failed — Anthropic apparently requires header/fingerprint details only
# claude-agent already uses directly for code tasks — it was never routed through here. # the real CLI sends, which LiteLLM faithfully relays but a hand-built request won't
# have. Do NOT expect this to work for other callers (Hermes, generic HTTP clients) —
# they aren't the real CLI and can't reproduce that fingerprint.
- model_name: anthropic-claude
litellm_params:
model: anthropic/claude-sonnet-5
general_settings: general_settings:
forward_client_headers_to_llm_api: true
master_key: os.environ/LITELLM_MASTER_KEY master_key: os.environ/LITELLM_MASTER_KEY