This is the setup it grew into: five specialized agents, a dedicated non-root user, and a command center dashboard at agents.cloud.merox.dev. Everything was templated in the infra repo, so it could be rebuilt from scratch in 15 minutes.
I tore the whole agent stack down on 23 July 2026 (user, cron jobs, sudoers, systemd unit and dashboard) and removed it from the infrastructure repo. The files this post points to are preserved in the last commit that had them. Read it as a record of how it was built, and of what to do differently; the sudoers section below has the part I’d change first.
5
agents
0
tokens to collect data
~15 min
rebuild
Why split them at all#
One agent trying to do everything becomes a compromise. The system prompt grows, context gets polluted between domains, and the “personality” that makes an agent useful for one task makes it annoying for another.
| Agent | Purpose | Runs |
|---|---|---|
news | Morning briefing in Romanian: tech stack updates, CVEs filtered to installed stack, community news | Daily 04:00 UTC |
infra | K8s cluster + VPS health checks, security alerts | 2× daily (08:00 + 20:00 UTC) |
costs | Backup verification, resource tracking, storage trends | Weekly (Sun 09:00 UTC) |
dashboard | Nightly audit + improvement of the command center dashboard | Daily 23:00 UTC |
orchestrator | Monitors all agents, auto-fixes safe issues, proposes improvements via Telegram | Daily 12:00 UTC |
Each agent has its own workspace directory, its own AGENTS.md (operating instructions), SOUL.md (personality), and sometimes HEARTBEAT.md (what to do proactively). They send Telegram messages on their own, without being prompted.
Architecture#
flowchart TB phone["phone / laptop"] -- Telegram --> gw["openclaw gateway
tailnet:18789, user openclaw"] gw --> cc["Claude Code CLI
Pro OAuth"] cc --> news & infra & costs & dash["dashboard"] & orch["orchestrator"] news --> n1["web search
/srv/dashboard"] infra --> i1["kubectl · flux
talosctl · docker"] costs --> c1["S3 backups"] dash --> d1["/srv/dashboard
index.html"] orch --> o1["agents.json
proposals.json"]
The gateway runs as a systemd user service under a dedicated openclaw user. No root. The agent workspaces describe what each agent manages and how it should behave. Claude figures out the commands to run.
Security model#
This is where the previous setup was weak: running as root because it was easier. The new setup:
A user with the minimum it needs#
useradd -m -s /bin/bash openclaw
usermod -aG docker openclaw
loginctl enable-linger openclawSudoers, scoped to exact binaries#
Two files:
/etc/sudoers.d/openclaw, for infra tooling:
Defaults:openclaw !requiretty
openclaw ALL=(ALL) NOPASSWD: /usr/bin/kubectl
openclaw ALL=(ALL) NOPASSWD: /usr/bin/flux
openclaw ALL=(ALL) NOPASSWD: /usr/bin/talosctl
openclaw ALL=(ALL) NOPASSWD: /bin/systemctl status *
openclaw ALL=(ALL) NOPASSWD: /usr/bin/systemctl status *
openclaw ALL=(ALL) NOPASSWD: /usr/bin/node/etc/sudoers.d/openclaw-fix-perms, for permissions automation:
openclaw ALL=(root) NOPASSWD: /usr/local/bin/openclaw-fix-permssudo node runs any JavaScript as root, and sudo systemctl status opens a pager that can spawn a root shell. Both are root access with extra steps, which undoes the point of the file. Neither was needed: systemctl status works unprivileged (add the user to systemd-journal for the log lines), and nothing the agents did required Node as root. If you build this, leave both out.
Docker group membership handles container access. That is root-equivalent on the host too, since anyone who can start a container can mount /. kubectl and talosctl get dedicated config files copied to /home/openclaw/.kube/ and /home/openclaw/.talos/. The agents are explicitly blocked from reading age.key, *.sops.yaml, and .env files, and that is enforced in AGENTS.md rather than left to the model.
The Claude Code permissions problem#
The gateway runs as openclaw, but if you interact with workspace files as root (e.g. via Claude Code CLI), any file you touch becomes root:root and the agents can’t write to it. The fix is a script that runs every 5 minutes and corrects ownership:
# /usr/local/bin/openclaw-fix-perms
chown -R openclaw:openclaw /home/openclaw/.openclaw/
chown -R openclaw:openclaw /srv/dashboard/
# Update scripts stay root-owned intentionally
chown root:root /srv/dashboard/update-*.sh
chown -R openclaw:openclaw /srv/merox/src/content/blog/Root crontab: */5 * * * * /usr/local/bin/openclaw-fix-perms
Systemd service: ExecStartPost=/usr/bin/sudo /usr/local/bin/openclaw-fix-perms
This means: even if you edit workspace files as root, they’re back to openclaw ownership within 5 minutes. No manual fixup needed.
Gateway on the tailnet only#
{
gateway: {
bind: "tailnet",
auth: { mode: "token" }
}
}The gateway listens on the Tailscale interface directly: reachable from any Tailscale device without a port forward, but not exposed on the public internet.
Telegram allowlist#
{
channels: {
telegram: {
botToken: "...",
allowFrom: ["YOUR_NUMERIC_ID"]
}
}
}Only your Telegram user ID can interact with the bot. Anyone else gets ignored at the gateway level, not the agent level.
Workspace files#
OpenClaw’s workspace system is the key difference from a SKILL.md single-file approach. Each agent gets a directory with:
workspace-infra/
├── AGENTS.md # operating instructions — what to check, how to respond
├── SOUL.md # personality — paranoid SRE tone, factual, no false positives
├── HEARTBEAT.md # what to do proactively on each scheduled tick
└── TOOLS.md # local notes — paths, commands, how to update the dashboardThe content matters more than the format. Here’s what actually makes the infra agent useful:
AGENTS.md (excerpt):
## Security check (2x/day via heartbeat)
kubectl get nodes
kubectl get pods -A | grep -v Running | grep -v Completed
docker ps --format "{{.Names}}\t{{.Status}}" | grep -v "Up"
df -h
Report on Telegram ONLY if there is a real problem.
Do not send "everything is ok" every check.SOUL.md (infra agent):
You are an SRE with healthy paranoia. You don't dramatize, but you don't minimize.
- "Trust but verify" — check live, don't assume
- Silence is golden — no false positives; when you send, it's real
- When you don't know something, say soHEARTBEAT.md (news agent):
| |
The command center dashboard#
Each agent writes to /srv/dashboard/data/: JSON files that a static HTML page reads and renders:
| |
The page auto-refreshes every 60 seconds. Dark mode, card-based layout. No backend; nginx serves static files.

An nginx container handles the serving:
# /srv/docker/agents-dashboard/docker-compose.yml
services:
agents-dashboard:
image: nginx:alpine
volumes:
- /srv/dashboard:/usr/share/nginx/html:ro
networks:
network-cloud-merox:
ipv4_address: 172.25.10.90
labels:
- "traefik.enable=true"
- "traefik.http.routers.agents-dashboard.rule=Host(`agents.cloud.merox.dev`)"
- "traefik.http.routers.agents-dashboard.entrypoints=https"
- "traefik.http.routers.agents-dashboard.tls.certresolver=cloudflare"
- "traefik.http.routers.agents-dashboard.middlewares=middlewares-authentik@file,default-headers@file"Protected by Authentik, the same SSO as the rest of the homelab stack.
After each run, an agent updates its status in agents.json:
import json
with open('/srv/dashboard/data/agents.json') as f:
data = json.load(f)
data['infra'] = {
'lastRun': '2026-05-28T08:00:00Z',
'status': 'ok', # ok / warn / error
'summary': 'All nodes healthy. Disk at 45%.'
}
with open('/srv/dashboard/data/agents.json', 'w') as f:
json.dump(data, f, indent=2)OpenClaw config#
All agents in one openclaw.json:
| |
OpenClaw uses Claude Code CLI’s OAuth, so there is no separate Anthropic API key if you have Claude Pro. All messages go to the default agent (news), which routes to the others via agentToAgent when needed.
Setup#
Full install in 15 minutes from scratch:
1. Install OpenClaw#
curl -fsSL https://deb.nodesource.com/setup_24.x | sudo -E bash -
sudo apt install -y nodejs
sudo npm install -g openclaw@latest
openclaw --version2. Create the service user#
sudo useradd -m -s /bin/bash -d /home/openclaw -c "OpenClaw Service" openclaw
sudo usermod -aG docker openclaw
sudo loginctl enable-linger openclaw
# Copy sudoers from infra repo
sudo cp /srv/kubernetes/infrastructure/agent/scripts/sudoers-openclaw /etc/sudoers.d/openclaw
sudo chmod 440 /etc/sudoers.d/openclaw3. Kubeconfig and talosconfig#
sudo -u openclaw mkdir -p /home/openclaw/.kube /home/openclaw/.talos
sudo cp /srv/kubernetes/infrastructure/kubeconfig /home/openclaw/.kube/config
sudo cp /srv/kubernetes/infrastructure/talos/clusterconfig/talosconfig /home/openclaw/.talos/config
sudo chown openclaw:openclaw /home/openclaw/.kube/config /home/openclaw/.talos/config
sudo chmod 600 /home/openclaw/.kube/config /home/openclaw/.talos/config4. Configure and install the workspaces#
sudo -u openclaw mkdir -p /home/openclaw/.openclaw
# Copy config (fill in Telegram token + your user ID)
sudo cp /srv/kubernetes/infrastructure/agent/openclaw.json.example \
/home/openclaw/.openclaw/openclaw.json
sudo chown openclaw:openclaw /home/openclaw/.openclaw/openclaw.json
sudo chmod 600 /home/openclaw/.openclaw/openclaw.json
# Install workspaces
AGENT_DIR=/srv/kubernetes/infrastructure/agent/workspaces
for agent in news infra costs dashboard orchestrator; do
sudo -u openclaw cp -r $AGENT_DIR/$agent \
/home/openclaw/.openclaw/workspace$([ "$agent" = "news" ] && echo "" || echo "-$agent")
done
sudo -u openclaw mkdir -p /home/openclaw/.openclaw/workspace/memory5. Authenticate Claude Code, then OpenClaw#
# 1. Log in to Claude Pro via OAuth (opens browser link)
sudo -u openclaw claude login
# 2. Run OpenClaw onboard to wire up Claude CLI auth and generate the gateway token
# This adds agentRuntime: { id: "claude-cli" } to openclaw.json — no API key needed
sudo -u openclaw XDG_RUNTIME_DIR=/run/user/$(id -u openclaw) \
openclaw onboard --non-interactive \
--mode local \
--auth-choice anthropic-cli \
--skip-bootstrap \
--skip-skills \
--skip-daemon \
--accept-riskThis uses your Claude Pro subscription: no per-token billing and no separate API key. The onboard command generates gateway.auth.token and adds agentRuntime: { id: "claude-cli" } for all Claude models. After this, re-add your Telegram credentials to ~/.openclaw/openclaw.json if onboard overwrote them.
6. The dashboard#
| |
7. fix-perms and the crontab#
| |
8. Start the gateway#
sudo -u openclaw mkdir -p /home/openclaw/.config/systemd/user
sudo cp /srv/kubernetes/infrastructure/agent/scripts/openclaw-gateway.service \
/home/openclaw/.config/systemd/user/
sudo chown openclaw:openclaw /home/openclaw/.config/systemd/user/openclaw-gateway.service
XDG_RUNTIME_DIR=/run/user/$(id -u openclaw) \
sudo -u openclaw systemctl --user daemon-reload
XDG_RUNTIME_DIR=/run/user/$(id -u openclaw) \
sudo -u openclaw systemctl --user enable --now openclaw-gateway.serviceVerify:
XDG_RUNTIME_DIR=/run/user/$(id -u openclaw) \
sudo -u openclaw systemctl --user status openclaw-gateway
sudo -u openclaw openclaw doctorProactive agents#
Five agents run without being asked:
infra at 08:00 and 20:00 UTC checks cluster nodes, unhealthy pods, disk space, stopped containers. Sends Telegram only if something is wrong. If you stop getting messages, either everything is fine or the agent itself died (check the service).
news at 04:00 UTC reads pre-fetched GitHub release data, then searches Hacker News, Reddit r/homelab and r/selfhosted, and does targeted web searches for CVEs and AI/infra news. CVEs are filtered to what’s actually installed on the K8s cluster and Oracle VPS, with no generic “this library is used somewhere” noise. Writes the structured JSON for the dashboard, the HTML briefing page, then sends a Telegram summary in Romanian.
costs on Sunday at 09:00 UTC verifies backups: Garage S3 object counts, Longhorn snapshot status, NAS accessibility. Writes backup.json for the dashboard and sends a summary on Telegram if anything is stale or missing.
dashboard at 23:00 UTC audits the command center dashboard nightly: validates JSON data files, checks that all JavaScript references have matching HTML elements, then makes one incremental improvement if the audit passes. Self-tests after each change.
orchestrator at 12:00 UTC audits all other agents: checks they ran on schedule, rotates oversized logs, validates data files. If it finds a pattern worth fixing in an agent’s logic, it sends a Telegram proposal with a concrete description. You approve or reject with /da or /nu. Applied changes are backed up automatically with rollback detection.
The orchestrator can also propose infrastructure-level changes (crontab schedule adjustments, new workspace files, openclaw.json agent additions), all requiring explicit approval before any change is applied. It never modifies security-sensitive config (auth tokens, allowlists, sudoers).
Heartbeats are configured via cron jobs on the openclaw user, not via OpenClaw’s built-in heartbeat polling (which burns tokens every 30 minutes even when there’s nothing to do):
| |
Orchestrator proposals#
The orchestrator introduces an approval loop for agent improvements. When it detects a pattern worth fixing (say, the news agent has been alerting on releases that Renovate would handle anyway, three weeks in a row), it writes a proposal to /srv/dashboard/data/proposals.json and sends a Telegram message:
Propunere #prop-20260529-news-001
Agent: news
Risc: low
Skip Renovate-covered releases in news alerts
News agent has alerted on 4 simple version bumps this week that Renovate
would have caught on Saturday. Add a filter for releases where no CVE
or breaking change is present.
Răspunde cu /da prop-20260529-news-001 sau /nu prop-20260529-news-001Reply /da and the next orchestrator run patches the relevant section of that agent’s AGENTS.md, backs up the original, and confirms on Telegram. If the change causes a regression (detected by comparing agents.json status before/after), it auto-rolls back and alerts.
Safe fixes like log rotation, missing JSON keys and stale file cleanup happen automatically without asking.
Disaster recovery#
Everything except secrets and memory was in the infra repo at agent/. Rebuilding on a new server took ~15 minutes:
- Install Node.js 24 + OpenClaw (
npm install -g openclaw@latest) - Create
openclawuser + docker group + linger - Copy
sudoers-openclaw→/etc/sudoers.d/openclaw - Copy kubeconfig + talosconfig to
/home/openclaw/.kube/and/home/openclaw/.talos/ - Install talosctl to
/usr/local/bin/(mise installs it per-user, not system-wide) - Copy all 5 workspaces from
agent/workspaces/and creatememory/dirs for all - Copy
openclaw.json.example→~/.openclaw/openclaw.json, fill in Telegram token sudo -u openclaw claude login(OAuth, opens a browser)sudo -u openclaw openclaw onboard --auth-choice anthropic-cli --skip-bootstrap- Install
openclaw-fix-permsto/usr/local/bin/+ root crontab*/5 * * * * - Install
openclaw-crontabas openclaw user’s crontab - Copy
agent/dashboard/scripts/*.sh→/srv/dashboard/, create/srv/dashboard/.envfrom template - Start systemd user service +
docker compose up -dfor dashboard
Can’t be auto-recovered:
openclaw.jsonsecrets (Telegram token + gateway auth token): keep in a password manager/srv/dashboard/.env(Telegram + Garage S3 + Apple credentials): keep in a password managerworkspace*/memory/files: agents rebuild context over a few days, history is lost
What the agents can’t do without attention on first boot:
infra-extended.jsonwon’t exist until the infra agent runs (08:00 UTC)proposals.jsonneeds to be initialized:echo '{"pending":[],"history":[]}' > /srv/dashboard/data/proposals.json
Everything else is in git or regenerates automatically.
The design decision worth stealing from this, if you take nothing else: the agents only speak when something is wrong, and everything that can be collected by a bash script is collected by a bash script. Built-in heartbeat polling burns tokens every thirty minutes to discover that nothing happened. Cron does the same job for free, and an agent that messages you daily regardless is one you stop reading by the second week.
Config template, workspace files, sudoers and systemd unit, as they were before the teardown: meroxdotdev/infrastructure, agent/.
