A lightweight Telegram bot that acts as a personal AI assistant. Written in Go — runs on a NAS, Raspberry Pi, or any small server. Bring your own API keys, no subscriptions.
- Multi-model routing — any configured model can be primary; automatic fallback on errors or rate limits; dedicated reasoner for complex tasks; Gemini-only multimodal role for images/audio; classifier-based routing to the reasoner
- Virtual slots — one slot per routing role (simple/default/complex/classifier/compaction/fallback/multimodal). Each slot can be backed by any provider type (OpenRouter, Gemini direct, or Claude Bridge) and any model from that provider's catalog. Cross-provider swaps happen at runtime via the admin UI, no restart.
- Vision- and tool-aware routing — model capabilities (vision, tool calling, reasoning, prices, context length) are fetched at startup from both OpenRouter's
/api/v1/modelsand Google's/v1beta/models, cached in the database, and used by the router to pick the right model per request. Multimodal role is UI-restricted to Gemini vision models. - Runtime model + provider swaps — change the provider type AND model id backing any role without restarting or editing YAML. Selections persist in
kv_settings.routing.overrides.SlotOverridesand survive restarts;BackendFactoryrebuilds the provider instance when the type changes. - Admin UI — htmx-driven web dashboard with five tabs: Routing & Models (per-role provider+model, OR/Gemini catalog browser with Pareto-frontier Suggest scored by AA benchmarks), Analytics (usage & cost), Prompts (system + classifier prompt overrides), Settings (scalar config — classifier_timeout, tool_filter_top_k, feature flags, tts.voice, etc.), and MCP (add/edit/delete MCP servers; auto-mirrors to Claude Bridge workspace when
MCP_BRIDGE_EXPORT_PATHis set). - Usage & cost analytics — every LLM call logged to a
usage_logtable with token counts (incl. cached prompt / reasoning splits), latency, and authoritative USD cost from OpenRouter'susage.cost/ Claude Bridge'stotal_cost_usd; admin dashboard at/analyticsshows totals cards with monthly forecast, daily stacked charts, top models by cost, role breakdown, cache hit rate, reasoning overhead, and "expensive turns" with question/answer JOIN viamessagesFKs. - Claude via bridge — use Claude (Anthropic Max subscription) as a provider through a lightweight host-side bridge service that wraps
claude -pCLI; no separate API key needed - Local Ollama — native
/api/chatprovider; still usable as an inference slot via the admin UI, but default production deploys route the classifier through Gemini/OpenRouter (Ollama stays available for embedding workloads and experimental local models) - Voice messages — send a voice message in Telegram and it's automatically transcribed via the multimodal model (Gemini), then processed as text through the normal pipeline; replies include both text and a voice message via Edge TTS
- Voice API + Atom Echo — HTTP and WebSocket voice API for hardware voice assistants; ships with ESPHome firmware for M5Stack Atom Echo (push-to-talk, LED feedback, streaming audio over WebSocket)
- Text-to-speech — Edge TTS (Microsoft) integration: no API key, high quality Russian/English voices, MP3 output for Telegram, WAV for hardware devices
- Web search — built-in
web_searchtool with provider dispatch: Tavily (free 1000/mo, LLM-ready answer + sources) or legacy Ollama Cloud; embedding-based query cache dedupes repeats across 6 hours - Web fetch — built-in
web_fetchtool that extracts main article text from a URL; HTTP + go-readability primary path, headless Chrome (CDP) fallback for JS-heavy or bot-protected pages - Filesystem tools — built-in
fs_list,fs_read,fs_write,fs_append,fs_delete,fs_searchscoped to a configurable directory; path traversal protection; great for personal notes, reference materials, and shared context with Claude Bridge - Semantic conversation memory — user messages are embedded and stored; within a session, relevant past turns are retrieved by cosine similarity instead of just "last N messages"
- Cross-session memory — past conversations are searched across all sessions; relevant snippets are automatically injected into the system prompt so the bot remembers what you discussed weeks ago
- Image support — send a photo (with or without caption) and it's routed automatically to the vision model
- Reply context — replying to a bot message or your own message prepends the quoted text so the LLM has full context
- Forwarded messages — forward any message (text, photo, link) to the bot, then ask your question; messages arriving within 2 s are batched automatically; when embeddings are configured and more than 3 messages are buffered, only the forwards most relevant to your question (cosine ≥ 0.25) are included — irrelevant topic noise is filtered out automatically
- Link extraction — hidden hyperlinks (
text_linkentities) in forwarded messages are surfaced as plain URLs for the LLM - MCP tool support — connects to any MCP-compatible server (HTTP/SSE), same
mcp.jsonformat as Claude Desktop; per-serverallowTools/denyToolsfiltering; vector similarity filtering selects only the most relevant tools per request - Configurable embeddings — shared embedding layer used for both tool filtering and conversation memory; supports Gemini (default), HuggingFace TEI, or any OpenAI-compatible endpoint
- Persistent memory — PostgreSQL-backed conversation history with automatic session management (SQLite fallback for local dev)
- Token-based compaction — auto-summarises old history when estimated token count exceeds threshold; images count as 1000 tokens each
- Smart token usage — classifier input truncated to 500 chars; large tool results (>2 KB) auto-summarised before entering history; response cache with cosine ≥ 0.92 threshold and 4-hour TTL
- Rich formatting — Markdown converted to Telegram HTML; responses ≥ 4096 chars sent as
response.md - Access control — allowlist by chat ID + owner-only enforcement
- Date/time awareness — current date and time injected into every request; timezone set via
TZenv var - Conversation stats —
/statsshows message count, character count, last compaction, and last activity
- Go 1.26+ (or Docker)
- Telegram Bot Token
- At least one LLM API key. Recommended: OpenRouter (one key, hundreds of models) + Gemini (direct, for vision/transcription fallback). Optional: Tavily for web search (1000 free searches/month, no card).
git clone https://github.com/dzarlax/personal_assistant.git
cd personal_assistant
# Setup configs from examples, create data dir
./scripts/setup.sh
# Fill in your API keys and Telegram token
nano .env
# Start
make docker-up
make logsconfig.yaml is baked into the Docker image — production deploys only need
docker-compose.yml and .env. Per-deployment overrides (routing, prompts,
MCP servers, scalar settings) live in the database and are edited through
the admin UI after first boot.
mkdir -p ~/personal_assistant/{data,bridge}
cd ~/personal_assistant
# Download docker-compose and the .env template
REPO="https://raw.githubusercontent.com/dzarlax/personal_assistant/main"
curl -sLO "$REPO/docker-compose.yml"
curl -sL "$REPO/.env.example" -o .env
curl -sL "$REPO/bridge/update.sh" -o bridge/update.sh && chmod +x bridge/update.sh
# Fill in your API keys and Telegram token
nano .env
# Start
docker compose up -dAfter the first boot, open the admin UI (/) and use the Routing & Models, Prompts, Settings, and MCP tabs to customise your deployment without touching files.
Get your Telegram chat ID: send /start to @userinfobot.
Use Claude (Anthropic Max/Pro subscription) as an LLM provider via a host-side bridge service. Activated on demand via /claude command in the bot.
Prerequisites: Claude Code CLI installed and logged in on the host.
# 1. Run setup (creates context dir, builds/downloads bridge, generates token)
./scripts/setup.sh --with-claude /path/to/assistant_context
# 2. Edit .env — fill in API keys, verify CLAUDE_BRIDGE_TOKEN and PROJECT_DIR.
nano .env
# 3. Edit config/mcp.json — add your MCP servers with "type": "http" field.
# Required for Claude Code to recognize HTTP-based MCP servers.
nano config/mcp.json
# 4. Install systemd service (production)
sudo cp bridge/claude-bridge.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now claude-bridge
# 5. Start bot
make docker-upA pre-built binary is published to GitHub Releases on every push to main that changes bridge/. To update:
./bridge/update.shThis downloads the latest binary and restarts the systemd service.
make run # local (loads .env, runs go)
make docker-up # Docker (copies missing example configs, starts)
make logs # follow Docker logsData is stored in PostgreSQL (set DATABASE_URL in .env). Falls back to ./data/conversations.db (SQLite) if no database URL is configured.
.env # secrets — not in git
.env.example # template
config/
config.yaml # model and routing config
system_prompt.md # fallback system prompt (used when filesystem is disabled)
mcp.json # MCP servers — not in git
mcp.json.example # template
routing.json # runtime routing overrides — auto-created
data/ # SQLite fallback DB — not in git
bridge/ # claude-bridge host service (optional)
main.go # HTTP → claude -p wrapper (config via env vars)
firmware/
atom-echo/ # ESPHome firmware for M5Stack Atom Echo
atom-echo.yaml # ESPHome config
secrets.yaml # WiFi + API token — not in git
components/ # custom voice_client ESPHome component
scripts/
init-context.sh # creates assistant_context directory
templates/
CLAUDE.md # system context template for Claude sessions
settings.json # permissions template for Claude CLI
All values support ${ENV_VAR} substitution. models: is a free-form map: each entry's key is the name referenced from routing.*, and its provider: field selects the backend (openrouter, gemini, ollama, claude-bridge, local, hf-tei, openai). The special key embedding is reserved for the MCP/memory embedding provider.
telegram:
bot_token: ${TELEGRAM_BOT_TOKEN}
allowed_chat_ids:
- ${TELEGRAM_OWNER_CHAT_ID}
owner_chat_id: ${TELEGRAM_OWNER_CHAT_ID}
models:
# --- OpenRouter slots — one per routing role so the admin UI can reassign
# each independently without affecting the others. The initial model
# ids below are starting points; change them via the web admin's
# "Suggest" button (filters the catalog by role-appropriate criteria). ---
simple-or:
provider: openrouter
model: qwen/qwen3.5-flash-02-23 # V T R, 1M ctx, $0.065/$0.26
api_key: ${OPENROUTER_API_KEY}
max_tokens: 2048
base_url: https://openrouter.ai/api/v1
default-or:
provider: openrouter
model: qwen/qwen3.5-flash-02-23 # same initial; diverge in UI
api_key: ${OPENROUTER_API_KEY}
max_tokens: 4096
base_url: https://openrouter.ai/api/v1
complex-or: # OR reasoner for when no bridge
provider: openrouter
model: qwen/qwen3-235b-a22b-thinking-2507 # thinking variant, $0.13/$0.60
api_key: ${OPENROUTER_API_KEY}
max_tokens: 8192
base_url: https://openrouter.ai/api/v1
compaction-or:
provider: openrouter
model: qwen/qwen3.5-9b # 262k ctx, $0.10/$0.15
api_key: ${OPENROUTER_API_KEY}
max_tokens: 2048
base_url: https://openrouter.ai/api/v1
# --- Gemini (direct Google API) — fallback + multimodal/voice-transcription. ---
gemini-flash-lite:
provider: gemini
model: gemini-3.1-flash-lite-preview
api_key: ${GEMINI_API_KEY}
max_tokens: 2048
base_url: https://generativelanguage.googleapis.com/v1beta/openai/
gemini-flash:
provider: gemini
model: gemini-3-flash-preview
api_key: ${GEMINI_API_KEY}
max_tokens: 4096
base_url: https://generativelanguage.googleapis.com/v1beta/openai/
# --- MCP / memory embedding. ---
embedding:
provider: hf-tei # or "openai" / "gemini"
base_url: https://embed.yourdomain.com
api_key: ${EMBED_API_KEY} # hf-tei: "user:password" for Basic Auth; openai: API key
# --- Local Ollama — cheap/fast classifier on host GPU. ---
classifier:
provider: ollama
model: qwen3:0.6b
base_url: http://host.docker.internal:11434
max_tokens: 64
no_think: true
# --- Claude via host-side bridge (see Claude Bridge section). Optional. ---
# claude-bridge:
# provider: claude-bridge
# base_url: http://host.docker.internal:9900
# api_key: ${CLAUDE_BRIDGE_TOKEN}
# max_tokens: 120 # CLI timeout in seconds
# --- Local llama.cpp server (OpenAI-compatible). Optional. ---
# local-gemma:
# provider: local
# model: gemma-3n-E2B
# max_tokens: 64
# base_url: http://classifier:8080/v1
routing:
simple: simple-or # level 1: simple/cheap tasks
default: default-or # level 2: moderate tasks (agentic loop)
complex: claude-bridge # level 3: Claude via bridge; swap to complex-or without bridge
compaction: compaction-or
fallback: gemini-flash-lite # direct Google — different vendor survives OR outage
multimodal: gemini-flash # direct Google — native input_audio for voice
classifier: classifier # local Ollama rates complexity 1/2/3
classifier_min_length: 0 # 0 = always; >0 = min chars; <0 = disabled
tool_filter:
top_k: 20 # top-K MCP tools selected per request via vector similarity; 0 = disabled
# Web search — provider: "tavily" (free 1000/mo, LLM-ready answer+sources) or "ollama".
web_search:
enabled: true
provider: tavily
api_key: ${TAVILY_API_KEY}
# Web fetch — extract main article text from a URL.
# cdp_url is optional; when set, JS-heavy / bot-protected pages fall back to a headless Chrome.
web_fetch:
enabled: true
cdp_url: ${WEB_FETCH_CDP_URL:-} # e.g. http://infra-chrome:9222
# Filesystem access — gives the assistant read/write access to a local directory
# filesystem:
# enabled: true
# root: /assistant_context # mount via docker-compose volumesModel capabilities: at startup the bot calls OpenRouter's /api/v1/models once, upserts every entry into the model_capabilities table (SQLite in dev, Postgres in prod), and applies each slot's caps so vision-aware routing knows which OpenRouter model can actually take images. When the admin UI (coming in Phase 2) swaps a slot's model, its caps are looked up from the store and SetProviderModel persists the choice in kv_settings under routing.overrides — no YAML edit or restart required.
Same format as Claude Desktop. Supports custom headers for auth and per-server tool filtering.
{
"mcpServers": {
"my-server": {
"url": "${SERVER_URL}",
"headers": {
"Authorization": "Bearer ${TOKEN}"
},
"denyTools": ["dangerous_tool"],
"allowTools": []
}
}
}denyTools— block specific tools, allow the restallowTools— allow only listed tools, block the rest- Omit both to allow all tools from the server
Plain text or Markdown injected as system prompt on every request.
When filesystem.enabled is true and a CLAUDE.md file exists at filesystem.root, the bot reads it as the system prompt instead. This allows sharing a single prompt file between the bot and Claude Bridge — no duplication, no sync issues.
When enabled, the bot gains built-in file management tools scoped to the configured root directory:
| Tool | Description |
|---|---|
fs_list |
List files and directories |
fs_read |
Read file contents (max 512 KB) |
fs_write |
Create or overwrite a file |
fs_append |
Append to a file (creates if missing) |
fs_delete |
Delete a file |
fs_search |
Case-insensitive text search across files |
All paths are relative to the root directory. Path traversal (../) is blocked. Mount the directory into the container via docker-compose volumes:
volumes:
- /path/to/context:/assistant_context:rwText-to-speech via Edge TTS (Microsoft). No API key required. Used for Telegram voice replies and Atom Echo responses.
tts:
enabled: true
voice: ru-RU-DmitryNeural # also: en-US-EmmaMultilingualNeural, uk-UA-OstapNeural
# rate: "+0%" # speech speed adjustment
# pitch: "+0Hz" # pitch adjustmentWhen enabled: voice messages in Telegram get a voice reply alongside text. The Voice API uses TTS for all responses.
HTTP + WebSocket voice API for hardware voice assistants (Atom Echo, etc).
voice_api:
enabled: true
listen: ":8086"
token: ${VOICE_API_TOKEN} # Bearer auth
chat_id: 9999 # separate conversation history from TelegramEndpoints:
POST /voice— send WAV audio, receive MP3 or WAV response (setAccept: audio/wavfor hardware devices)GET /voice/ws?token=<token>— WebSocket for streaming: send binary PCM frames, receive binary WAV framesGET /voice/health— health check
WebSocket protocol:
- Client connects to
/voice/ws?token=<token> - Client sends binary frames with raw PCM audio (16kHz 16bit mono)
- Client sends text frame
{"action":"stop"}to end recording - Server sends text frame
{"status":"processing","transcription":"..."}after STT - Server sends binary frames with WAV audio response
- Server sends text frame
{"status":"done","response":"..."}when complete
Exposed via Traefik at voice.dzarlax.dev.
ESPHome firmware for M5Stack Atom Echo in firmware/atom-echo/:
cd firmware/atom-echo
cp secrets.yaml.example secrets.yaml # fill in WiFi + API token
esphome run atom-echo.yaml # compile + flash via USBFeatures:
- Push-to-talk button → stream audio via WebSocket → play response
- LED feedback: blue (idle) → orange (connecting) → red (recording) → yellow pulse (processing) → green (playing)
- 16kHz 16bit mono mic with 4x gain
- Streaming playback (no buffer limit for responses)
- Up to 10s recording via WebSocket streaming
| Command | Description |
|---|---|
/clear |
Reset conversation context |
/compact |
Summarise and compress history manually |
/stats |
Show conversation stats (messages, chars, last compaction) |
/model |
Show current model |
/model list |
List all available models |
/model <name> |
Switch to a specific model for the session (e.g. /model deepseek-r1) |
/model reset |
Back to auto-routing |
/claude <question> |
Enter Claude mode — sends question via Claude Bridge |
/exit |
Exit Claude mode, back to auto-routing |
/routing |
Configure routing roles permanently via inline keyboard. Appends an "Open Admin UI" button when admin_api.base_url is set. |
/tools |
List connected MCP tools grouped by server |
/help |
Show help |
Note:
/modelis a temporary session override — it resets on restart. To permanently change the primary model use/routingor the Admin UI.
Optional web interface for browsing OpenRouter's full model catalog, filtering by price / capabilities (Free / Vision / Tools / Reasoning), and assigning any model to a slot with a click. Also edits routing roles (default, reasoner, multimodal, fallback, classifier).
Enable by setting admin_api.enabled: true in config.yaml and providing a token or forward-auth setup:
admin_api:
enabled: true
listen: ":8087"
token: ${ADMIN_API_TOKEN}
trust_forward_auth: true # when behind Traefik/Authentik
forward_auth_header: X-authentik-username
base_url: https://assistant.example.com # surfaced in Telegram /routingStack: Go stdlib HTTP + htmx + dzarlax design system, everything embedded into the binary via go:embed. Static assets (CSS + JS) are re-downloaded on every CI build so the design system stays fresh (ARG ASSETS_CACHEBUST in Dockerfile).
Auth precedence:
- Authentik forward-auth — when
trust_forward_auth: trueand the upstream middleware setsX-authentik-username(or the configured header), the request is authenticated. - Cookie
admin_auth— set automatically after a bootstrap?token=<ADMIN_API_TOKEN>visit. - Bearer token —
Authorization: Bearer <ADMIN_API_TOKEN>forcurl/ monitoring. - Otherwise 401.
Deployment recipe (behind Traefik + Authentik):
# docker-compose.yml
labels:
- "traefik.enable=true"
# Telegram Mini App: public at the proxy layer, protected by Telegram initData
# HMAC + TELEGRAM_OWNER_CHAT_ID inside the Go service.
- "traefik.http.routers.assistant-tg-admin.entrypoints=https"
- "traefik.http.routers.assistant-tg-admin.rule=Host(`assistant.example.com`) && PathPrefix(`/tg-admin`)"
- "traefik.http.routers.assistant-tg-admin.priority=100"
- "traefik.http.routers.assistant-tg-admin.tls=true"
- "traefik.http.routers.assistant-tg-admin.tls.certresolver=letsEncrypt"
- "traefik.http.routers.assistant-tg-admin.service=assistant-admin"
# Full web admin: keep Authentik/forward-auth protection.
- "traefik.http.routers.assistant-admin.entrypoints=https"
- "traefik.http.routers.assistant-admin.rule=Host(`assistant.example.com`)"
- "traefik.http.routers.assistant-admin.priority=10"
- "traefik.http.routers.assistant-admin.tls=true"
- "traefik.http.routers.assistant-admin.tls.certresolver=letsEncrypt"
- "traefik.http.routers.assistant-admin.middlewares=authentik-auth"
- "traefik.http.services.assistant-admin.loadbalancer.server.port=8087"Pair with an Authentik Application + Provider (Proxy) configured to protect the route; restrict access via policy binding to a admins group or similar.
Every incoming message is routed to one of seven roles. Roles are named after what they serve, so the same name maps consistently through config.yaml → routing.*, Go code, and the admin UI.
| Role | Level | When it fires | Typical model |
|---|---|---|---|
simple |
L1 | Classifier rates the message as trivial (greeting, chitchat, one-line lookup). Cost-optimised. | A cheap OpenRouter model, e.g. google/gemini-2.5-flash-lite |
default |
L2 | Most messages. The agentic loop (tool calling, memory, MCP) lives here. | A mid-tier OpenRouter model, e.g. deepseek/deepseek-chat-v3.1 |
complex |
L3 | Classifier rates the message as hard (deep reasoning, proofs, hard debugging), or /model <name> override. |
A strong model, e.g. Claude Sonnet 4.5 via bridge or OpenRouter |
fallback |
— | default provider is unavailable (5xx / 429 / network). Intentionally a different vendor from default for provider-outage resilience. |
gemini-flash-lite (direct Google) |
multimodal |
— | Message contains an image or audio (voice transcription). default/simple are preferred first if they support vision; multimodal is used otherwise. |
Gemini Flash (direct Google, for native input_audio support) |
classifier |
— | Rates user text as 1/2/3 to pick simple/default/complex. Returns one digit — no tools, no history. |
Local Ollama (e.g. qwen3:0.6b) or a free OpenRouter model |
compaction |
— | Summarises old history when the conversation grows. No tools. | Anything cheap that understands long context |
The routing order for a normal request:
- Image or audio? →
multimodal(orsimple/defaultif they natively support vision). - Continuing a tool-calling loop? → Keep the same provider (stickiness, one model owns a tool chain from start to finish).
- Run
classifier, get 1 / 2 / 3 →simple/default/complex. defaultprovider returns 5xx/429/network? →fallback(same request, different vendor).
The classifier is a tiny call with no history and no tools — it outputs a single digit. It's cheap enough that classifier_min_length: 0 (always run) is fine when backed by a local Ollama model. Raise classifier_min_length to skip trivially short messages if you run a paid classifier; set it negative to disable the three-level split entirely (everything goes through default).
All routing roles can be changed live from the Admin UI or from Telegram /routing. Changes persist across restarts in the database (kv_settings table, key routing.overrides). A legacy config/routing.json file is auto-imported and removed on first start, then never touched again. On startup, the bot notifies the owner via Telegram if any routing role references an unavailable model.
When an embedding model is configured, the bot gains two levels of long-term memory beyond the current session:
User messages are embedded and stored in the database. Instead of always taking the last 30 messages, the context window is built as:
- Last 10 messages — always included (recent context)
- Up to 20 older turns — selected by cosine similarity to the current query
Conversational turns (user message + assistant response + tool calls) are kept together to preserve coherence.
On each request, past sessions are searched for semantically similar conversations. Up to 5 relevant snippets (cosine similarity > 0.75) are injected into the system prompt:
---
Relevant context from previous conversations:
[2026-01-15] You: what did we decide about the database?
Assistant: We decided to stay with SQLite for simplicity...
Each snippet is truncated (200 chars user / 300 chars assistant) with a 3000-char total budget, so it never bloats the prompt.
This is complementary to the personal-memory MCP server: personal-memory stores explicit facts you choose to remember; cross-session RAG surfaces actual conversation fragments automatically.
- History persists across restarts (PostgreSQL; SQLite fallback)
- After 4 hours of inactivity, a new session starts automatically — the last summary is carried over
/cleardoes a full reset with no carry-over- Compaction triggers when estimated token count exceeds 16 000 tokens (images count as 1000 tokens each)
- When embeddings are configured, old messages are clustered by topic (cosine similarity < 0.65 starts a new cluster) and each cluster is summarised separately — producing a more structured, topic-aware summary. Falls back to single-pass summarisation when embeddings are unavailable
Use Claude from your Anthropic Max/Pro subscription as an LLM provider — no separate API key needed. A lightweight Go service runs on the host and wraps claude -p CLI. Activated on demand via /claude command.
/claude <question> → Bot (Docker) → POST /ask → claude-bridge (host:9900) → claude -p → response
- Bridge runs on the host (not in Docker) as a systemd service — it needs access to
claudeCLI - Bot in Docker reaches the bridge via
host.docker.internal:9900 - Bridge listens on the Docker bridge network (
172.17.0.1:9900) — not exposed externally - Each request: bot formats conversation history into a single text prompt →
claude -p→ response - Claude CLI reads
CLAUDE.mdand.mcp.jsonfrom the project context directory - MCP servers in
config/mcp.jsonmust have"type": "http"for Claude Code compatibility (bot ignores this field)
No source code on the server — only the binary, config, and docker-compose:
~/personal_assistant/
├── .env # secrets
├── docker-compose.yml # bot (Docker)
├── bridge/
│ ├── claude-bridge # binary (managed by systemd)
│ └── update.sh # update script
├── config/
│ ├── config.yaml # models and routing
│ ├── routing.json # runtime routing overrides
│ ├── mcp.json # MCP servers
│ └── system_prompt.md # fallback system prompt
└── data/
└── routing.json # runtime state
# Shared context directory (mounted into bot container as /assistant_context)
~/vol/assistant_context/
├── CLAUDE.md # unified system prompt (bot + Claude Bridge)
├── .mcp.json # MCP config for Claude CLI
├── notes/ # personal notes (read/write via filesystem tools)
├── reference/ # reference materials
└── tasks/ # task files and plans
# /etc/systemd/system/claude-bridge.service
[Unit]
Description=Claude Bridge - HTTP wrapper for Claude Code CLI
After=network.target
[Service]
Type=simple
ExecStart=/root/personal_assistant/bridge/claude-bridge
Environment=CLAUDE_BRIDGE_TOKEN=<your-token>
Environment=CLAUDE_BRIDGE_PROJECT_DIR=/root/vol/assistant_context
Environment=CLAUDE_BRIDGE_LISTEN=172.17.0.1:9900
Environment=CLAUDE_BRIDGE_CLI=/root/.local/bin/claude
Environment=CLAUDE_BRIDGE_CONCURRENCY=1
Environment=CLAUDE_BRIDGE_TIMEOUT=120
Environment=PATH=/root/.local/bin:/usr/local/bin:/usr/bin:/bin
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable --now claude-bridge~/personal_assistant/bridge/update.shDownloads the latest binary from GitHub Releases and restarts the service.
| Variable | Required | Default | Description |
|---|---|---|---|
CLAUDE_BRIDGE_TOKEN |
Yes | — | Shared secret (Bearer auth) |
CLAUDE_BRIDGE_PROJECT_DIR |
Yes | — | Path to project context directory |
CLAUDE_BRIDGE_LISTEN |
No | 127.0.0.1:9900 |
Set to 172.17.0.1:9900 for Docker access |
CLAUDE_BRIDGE_TIMEOUT |
No | 120 |
Default CLI timeout in seconds |
CLAUDE_BRIDGE_CONCURRENCY |
No | 1 |
Max parallel CLI calls |
CLAUDE_BRIDGE_CLI |
No | claude |
Path to CLI binary |
- ~5-8s per request (CLI cold start on each call)
- No streaming — user waits for full response
- No images — multimodal queries routed to Gemini
- Stateless — conversation history formatted into prompt by the bot
This bot is designed to work with self-hosted MCP servers. Two ready-made servers are available:
| Server | Description |
|---|---|
| personal-memory | Semantic long-term memory with vector embeddings + Todoist integration |
| health-dashboard | Health data from Apple Health (via Health Auto Export) with MCP tools for AI analysis |
Configure them in config/mcp.json.
flowchart TD
User(["📱 Telegram"])
Echo(["🔊 Atom Echo"])
VoiceAPI["Voice API\n(HTTP + WebSocket)\n:8086"]
TTS["Edge TTS\n(Microsoft)"]
Handler["Telegram Handler\n(debounce · reply chain · batch)"]
FwdBuf["Forward Buffer\n(embed · relevance filter)"]
Agent["Agent\nagentic loop · response cache"]
Router["LLM Router\n(primary · fallback · reasoner\nmultimodal · classifier)"]
MCP["MCP Client"]
Store[("PostgreSQL\n(SQLite fallback)")]
Emb["Embedding Model\n(Gemini / HF-TEI / OpenAI)"]
subgraph Memory ["Semantic Memory"]
SM["Within-session RAG\nlast 10 + top-20 by similarity"]
CM["Cross-session search\ntop-5 snippets → system prompt"]
SC["Semantic compaction\ntopic clusters → per-cluster summary"]
end
WebSearch["Web Search\n(Tavily / Ollama)"]
WebFetch["Web Fetch\n(readability + CDP fallback)"]
FS["Filesystem Tools\n(notes · reference · tasks)"]
CapStore[("model_capabilities\n+ kv_settings\n(persistent routing)")]
subgraph LLMs ["LLM Providers"]
OR["openrouter slots (simple-or, default-or, complex-or, compaction-or)"]
ORx["openrouter (complex-or, ...)"]
GM["gemini-flash-lite / gemini-flash"]
OLC["ollama (classifier, local)"]
CB["claude (via bridge, optional)"]
end
subgraph Host ["Host (outside Docker)"]
Bridge["claude-bridge\n:9900"]
CLI["claude -p CLI"]
CTX["assistant_context/\nCLAUDE.md · .mcp.json"]
end
subgraph Servers ["MCP Servers"]
S1["personal-memory"]
S2["health-dashboard"]
S3["..."]
end
Echo -->|"WSS audio stream"| VoiceAPI
VoiceAPI -->|"STT + Process + TTS"| Agent
VoiceAPI <-->|"MP3 → WAV"| TTS
Handler <-->|"voice reply"| TTS
User -->|"text / photo / voice / reply"| Handler
User -->|"forwarded messages"| FwdBuf
FwdBuf <-->|"embed forwards"| Emb
FwdBuf -->|"relevant forwards only\ncosine ≥ 0.25"| Handler
Handler --> Agent
Agent <-->|"store + retrieve"| Store
Store <-->|"embed messages"| Emb
Store --> SM
Store --> CM
Store --> SC
SM -->|"context window"| Agent
CM -->|"system prompt snippets"| Agent
Agent -->|"cache hit → skip LLM"| Agent
Agent --> Router
Agent -->|"query embed · top-K filter"| MCP
Agent <-->|"web_search tool"| WebSearch
Agent <-->|"web_fetch tool"| WebFetch
Agent <-->|"read · write · search"| FS
Router <-->|"caps + overrides"| CapStore
FS <-->|"shared directory"| CTX
Handler -->|"voice → transcribe"| Agent
MCP <-->|"embed tools at startup"| Emb
MCP <-->|"tools/call"| Servers
Router --> OR
Router --> ORx
Router --> GM
Router --> OLC
Router --> CB
CB -->|"POST /ask"| Bridge
Bridge -->|"claude -p"| CLI
CLI -->|"reads context"| CTX
See CLAUDE.md for developer details.
Removed direct integrations with DeepSeek, Qwen (DashScope), and Ollama Cloud. All cloud model access now goes through OpenRouter with a single API key. Local Ollama is retained for the classifier role. Gemini stays direct via Google's API for multimodal/vision (and voice transcription). Claude stays via the existing bridge.
Motivation:
- Ollama Cloud was consistently slow on the 397B model we were using — latency well above 10s to first token on normal-sized prompts, and response quality not noticeably better than cheaper alternatives.
- Managing 3–4 vendor accounts (DeepSeek + Qwen + Ollama Cloud + Gemini) for overlapping model access was friction: four billing pages, four API tokens, four SDK quirks to track. OpenRouter exposes DeepSeek, Qwen, Claude, Gemini, Llama, and 300+ other models behind one OpenAI-compatible endpoint and one key.
- Tool-call-aware fallbacks: OpenRouter's
provider.require_parameters=true+allow_fallbacks=trueguarantees the picked upstream supports tool calling (critical for the agentic loop) and transparently retries on a different upstream if the first returns an error. - Usage visibility:
usage.include=truereturns prompt/completion token counts on every response, making cost tracking straightforward. - Occasional free-tier promos — OpenRouter regularly has
:freevariants of popular models (e.g. Gemma, Llama) that are genuinely free, swappable at runtime.
Under the hood: ModelsConfig was refactored from a fixed Go struct to map[string]ModelConfig, so you can define as many OpenRouter slots as you want (each with a different model id) and assign each to a different routing role — no Go changes required. Model capabilities are fetched from OpenRouter's /api/v1/models at startup and cached in the database, making vision-aware routing accurate even when you swap the slot to a different model mid-session.