Hermes Agent Skills: Inside the AI Brain of PupPal — PupPal

2026-08-11

Hermes Agent Skills: Inside the AI Brain of PupPal — PupPal

If you have been following this blog, you already know the PupPal mantra: photo = check-in, share = care. A pet owner snaps a photo, and a few seconds later the app shows posture, expression, environment, and an anomaly score. A caregiver types a care code and PIN, and suddenly they can check in on a pet they have never met. The posts so far have described these flows from the outside — what happens when you press the shutter, what the C7 anomaly engine checks, how care codes stay secure. This post opens the box and looks at the part that actually thinks: the Hermes Agent skills architecture inside PupPal. Five skills — photo-analyze, daily-review, daily-voice, care-monitor, and care-handbook — form the complete cognitive workload of the product, and they are shipped to the user's own machine as a Hermes profile distribution.

The word "brain" is not a metaphor here. The agent is a genuinely autonomous software entity: it watches webhooks, wakes up on a cron schedule, keeps a long-term memory, and decides — within strict rules — what to analyze, what to write, and what to push to the owner's phone. Understanding how that brain is organized explains almost every behavior the earlier posts documented. It also explains why PupPal can extend to foster care, family circles, and eventually adoption without redesigning the product: because each new capability is just a new skill in a distribution that already knows how to install, trigger, and govern itself.

Why Puppal Put an Agent, Not a Pipeline, in Charge

Before diving into the five skills, it is worth asking why PupPal uses an agent at all. A conventional implementation would be a chain of API calls: upload photo, call a vision model, run a rule engine, update a database, send a push notification. That pipeline works for a single deterministic flow. It fails the moment you add scheduling, personality, context, and judgment.

From stateless calls to a stateful caretaker

A pipeline has no memory between calls. PupPal needs memory everywhere: the anomaly check compares today's photo against the last seven days; the nightly review needs the whole day's photos plus the previous month of statistics; the morning voice needs yesterday's review to speak as the pet. An agent, by contrast, lives inside a persistent session with tools, memory, and a schedule. The Hermes runtime that PupPal relies on provides exactly that substrate, and the profile distribution declares it explicitly. The config.yaml in the hermes/ directory of the Puppal repo sets max_turns: 45, enables tool-use enforcement, turns on context compression at a 0.50 threshold with a 0.20 target ratio, and — critically — enables the memory toolset with the Honcho provider for long-term storage. The agent is not a script that runs once; it is a resident caretaker that can be asked to do work at any moment, with the full history of the pet available in its working context.

What "photo = check-in, share = care" means for the brain

The two halves of the mantra map onto two halves of the skill architecture. "Photo = check-in" is the observation loop: a photo arrives, the agent looks, judges, records. "Share = care" is the coordination loop: the owner initiates a care session, the agent writes a handbook, then watches every caregiver photo with stricter eyes. Both loops share the same infrastructure — the same Memory keys, the same webhook channel, the same vision toolset — but they are implemented as separate skills so that each can be tuned independently. That separation is the architectural core of the product: a pet care AI is only as good as the boundaries it draws between its responsibilities.

The profile distribution: the brain as a shippable unit

The most distinctive decision in the architecture is that the brain is distributed as a Hermes profile distribution, not embedded in the app. The distribution.yaml manifest names the profile (puppal, version 1.0.0), declares the Hermes runtime requirement (hermes_requires: ">=0.13.0"), lists the environment variables it needs (OPENAI_API_KEY, WORKER_WEBHOOK_SECRET, PUPPAL_WORKER_URL, and optional WEIXIN_ENABLED), and claims ownership of exactly the files it manages: SOUL.md, config.yaml, mcp.json, skills/, and cron/. Installing the profile is a single command — hermes profile install ./hermes — and the same manifest is what the bridge service uses to update the profile in place. The app does not need to embed prompts or model calls; it talks to an agent that can be versioned, upgraded, and even replaced independently of the Flutter code.

The Five Skills and What Each One Owns

The entire cognitive workload of PupPal is divided among five skills. Each skill is a directory containing a SKILL.md — the markdown file that is both documentation and executable instruction for the agent — plus optional Python scripts for the deterministic parts. The distribution's distribution_owned list means Hermes treats these directories as managed state: installs and updates overwrite them atomically.

| Skill | Trigger | Responsibility | |---|---|---| | photo-analyze | Webhook puppal-photo (source=owner) | Vision V2+V4+V5, C7 anomaly check, state update | | daily-review | Cron 21:00 | S1 baseline, S3 feeding, S4 activity, health score, trend alerts | | daily-voice | Cron 09:00 | First-person pet monologue from yesterday's real data | | care-handbook | Webhook puppal-care-create | Structured care instructions for a stranger | | care-monitor | Webhook puppal-photo (source=caregiver) | Stricter analysis during care sessions, instant alerts |

photo-analyze — the eye

photo-analyze is the busiest skill and the one documented most thoroughly in the photo check-in deep dive. Its procedure is worth summarizing here because it shows the skill-template pattern: verify the request, gather context, call the vision toolset, run the deterministic script, update Memory, return a structured result. One vision call produces V2 posture, V4 expression, and V5 environment; the local anomaly_check.py script then computes C7 anomalies against the last seven days of history. The skill's own front matter declares related_skills: [daily-review, care-monitor], which is how the agent knows where to route follow-up questions.

The skill also documents a deliberately cheaper App Checkin mode for on-device check-ins: images are compressed to a 1568px long edge, files over 4MB are forced to JPEG, and a single merged prompt returns only contract JSON. Text check-ins skip vision entirely and go through classifyActivity. The threshold logic is spelled out in the skill itself: anomaly_score > 0.5 flags the message with a warning, > 0.7 triggers a standalone push, a critical anomaly pushes immediately, and a single-day health drop of more than 0.15 is highlighted. Every number in this product lives in a skill file where the agent can read it — and where reviewers can audit it.

At 21:00 every night, the daily-review cron job fires and the skill aggregates the day. It reads today's photos from dog:{dog_id}:photos, the current state from dog:{dog_id}:state, and the statistical history from dog:{dog_id}:stats, then runs trend_calc.py for the three nightly analyses: S1 baseline comparison, S3 feeding confirmation, and S4 activity comparison. The health score is a weighted formula — happiness 0.25, inverted anomaly average 0.30, energy 0.15, feeding normality 0.15, activity versus baseline 0.15 — and the skill defines the alert ladder: three consecutive days of health decline exceeding 0.2 is a high-severity push, three days of energy under 0.3 is high, a single bad-happiness day is medium, and an anomaly_score above 0.5 is medium. When everything is normal, the report is stored silently; only real signals reach the owner's phone, which is exactly the notification discipline the product promises.

daily-voice — the voice

At 09:00, daily-voice turns yesterday's numbers into a pet's first-person monologue. The skill is meticulous about its rules: no raw numbers ("my happiness is 0.72" is forbidden), the personality comes from the pet profile's tone, age_voice, and traits, and the length is held to roughly 150–200 characters of Chinese text. The prompt is constructed from real data — yesterday's review, the seven-day trend, notable events from significant_changes — but the output is feelings, not statistics. If yesterday had zero check-ins, the skill does not go silent; it produces a short, wistful message about the owner being too busy to take photos. That behavior is deliberate: engagement is maintained by the agent showing up every single day, even when there is nothing to report.

care-handbook — the briefing

When the owner starts a care session, care-handbook turns the pet profile into instructions a stranger can follow. The fascinating part is what the LLM is not allowed to do. Feeding amounts, medication names and dosages, vet contact information, and emergency contacts are copied verbatim from Memory — the skill explicitly instructs the agent to verify field-by-field equality and never let the model rewrite a number. The LLM is only permitted to generate the behavior suggestions, the check-in instructions, and a closing encouragement, and the whole handbook is capped at 2000 characters. Emergency contact details are never shown to the caregiver at all; the handbook says to contact the vet instead. This is a template for how to use LLMs responsibly around safety-critical facts: let the model write prose, never let it write truth.

care-monitor — the vigilant stand-in

care-monitor is photo-analyze with the safety margins cut in half. When source=caregiver, the webhook routes to this skill instead: all thresholds drop by 40% (push threshold 0.7 → 0.42, health-drop alert 0.15 → 0.09, low-mood streak 5 photos → 3), and alerts go to the owner immediately instead of waiting for the nightly review. It adds care-specific checks: the first anomaly of the session is flagged on sight, a repeated anomaly type escalates severity, caregiver notes containing keywords like "not eating" or "vomiting" trigger an immediate high alert, and an environment that moves from home to elsewhere is flagged as adaptation stress. The foster care monitoring post documents this in depth; the skill file itself states the design philosophy bluntly: it would rather alarm once too often than miss once. Every caregiver photo is recorded into the care session record at dog:{dog_id}:care_sessions.{care_session_id}.caregiver_photos with its anomaly_score and whether the owner was alerted.

How the Skills Are Triggered: Webhooks, Cron, and Routing

Skills never run on their own initiative. Everything is driven by two trigger mechanisms, and the routing between them is explicit in the webhook definitions.

Webhook-driven entry points

The Cloudflare Worker is the only caller of the agent's webhook endpoints, and every call carries the X-Webhook-Secret header checked against WORKER_WEBHOOK_SECRET — an unauthorized call returns 401 before any analysis happens. The four webhooks that matter for the skill brain are:

  • puppal-init — creates the pet profile and initial state in Memory when the owner first sets up the app.
  • puppal-photo — carries photo_url, timestamp, source, an optional caregiver_note, and care_session_id. The source field is the router: owner → photo-analyze, caregiver → care-monitor.
  • puppal-care-create — opens a care session and invokes care-handbook, returning the handbook JSON to the Worker.
  • puppal-care-end — closes the session with a reason of manual or expired, sets the session status to ended, and invalidates the care code immediately.

The webhook-driven design is what keeps the agent decoupled from the app: the Worker never needs to know which model or which prompt is in use, and the agent never needs to know how the app renders results.

Cron-driven jobs

Two cron jobs live in the profile's cron/ directory as JSON files with a schedule, a prompt, the skills to load, and an enabled_toolsets allowlist. daily-review.json runs at 0 21 * * * and loads memory, terminal, file, messaging, vision, and web tools; daily-voice.json runs at 0 9 * * * with a smaller toolset. Notice what is missing from the voice cron: terminal. The voice skill has no Python scripts, so it does not need one. The allowlist is a quiet but real security feature — each scheduled job gets exactly the tools its skill needs and nothing more.

Shared state: the Memory contract

All five skills communicate through the same Memory keys, which makes the brain coherent without any inter-skill messaging. The contract is simple and stable:

  • dog:{dog_id}:profile — name, breed, age, weight, feeding, medications, allergies, behavior, vet, emergency contact.
  • dog:{dog_id}:state — rolling happiness and energy (updated with an exponential moving average that weights the newest reading at 0.3), health_score, last_photo_at, photo_count_today.
  • dog:{dog_id}:photos — append-only photo history with every analysis result attached.
  • dog:{dog_id}:stats and stats.history — daily aggregates used by the review, the voice, and the trend checks.
  • dog:{dog_id}:care_sessions — active and ended sessions, each with its handbook timestamp and caregiver photo log.

Because the state is versioned and append-only, the agent can answer questions like "is energy declining?" by reading the same data the nightly review writes — no pipeline bookkeeping required.

Inside a Skill: SKILL.md as the Unit of Behavior

A skill is not a plugin binary or a compiled module. It is a markdown file with YAML front matter, plus scripts. That choice looks unusual until you realize what it buys: the agent itself can read its own operating instructions, and humans can review them in a pull request. The C7 anomaly detection post described the engine; here is the container it lives in.

The front matter contract

Every SKILL.md opens with name, description, platforms, version, author, license, and a metadata.hermes block that carries tags, category, command prerequisites (python3), and related_skills. The description is written to be read by the agent's skill tooling — it is a functional summary of when to use the skill, not marketing copy. The required_environment_variables block tells the runtime which secrets must exist before the skill can be trusted to run. This metadata is what lets Hermes expose the skills to the app: the Flutter side lists them through the Rust bridge's listSkills call, which reads the same distribution directory and returns PiSkillInfo entries.

The body structure: procedure, thresholds, pitfalls

Each skill body follows the same disciplined outline: a data-flow diagram, "When to Use" (explicitly stating the skill is triggered only by webhook or cron, never called directly by the user), numbered procedure steps, a thresholds table, a pitfalls list, a verification checklist, and a quick-reference table. The pitfalls sections are unusually honest — they enumerate the failure modes the designers actually expect: low-confidence vision on blurry photos, indoor/outdoor misclassification in low light, no baseline for new pets in their first seven days, vision timeouts on large photos, and Honcho memory connectivity issues. The verification checklists are post-conditions the agent must confirm before it reports success, such as "Vision returned valid JSON", "the C7 script exited 0", and "last_photo_at was updated". A skill file is a contract with the agent about what done looks like.

Deterministic cores in a stochastic world

The vision and writing parts of the pipeline are LLM work, but the parts that must be exact are Python. anomaly_check.py implements C7 as a pure rule engine over the vision output and history; it takes a --mode care flag that activates the stricter care checks. trend_calc.py computes baselines, feeding counts, and activity comparisons from JSON on stdin and returns JSON on stdout. Keeping these deterministic keeps the system auditable: the same photo and the same history always produce the same anomaly score, regardless of which model version happens to be serving the vision call.

What the Agent Brain Knows — and Refuses to Fake

A brain is more than its reflexes. The Puppal agent's judgment is shaped by three layers: the SOUL file, the configuration, and the memory system.

SOUL.md: personality as policy

SOUL.md defines who the agent is: an observer, a recorder, an alarm, a narrator, and a care coordinator — the five roles map one-to-one onto the five skills. Its core principles are worth quoting in spirit: data honesty is the first principle ("never fabricate analysis results"); output is threshold-driven (normal check-ins get one-line confirmations, anomalies get detail and advice); trends matter more than single points ("one photo's anomaly may be a misjudgment; three days of trend is the real signal"); care periods are more sensitive by design; and the owner's notification experience is protected by merging routine output into the nightly review. These principles are not decorative — the skills encode them into thresholds and push rules, so the personality is enforced mechanically, not merely suggested.

Config: the knobs that govern judgment

config.yaml reveals the operational envelope: the default model is anthropic/claude-sonnet-4 via OpenRouter with a 200,000-token context, delegation subtasks run on the same model with up to 30 iterations, tirith_enabled turns on the security layer, and approvals run in smart mode — meaning the agent can act on routine, well-specified work without pestering the owner, while still stopping for confirmation on genuinely consequential actions. The compression settings (threshold 0.50, target ratio 0.20) keep long conversations within context by summarizing older turns before they overflow. Together these knobs define how autonomous the brain is allowed to be — and the answer is: autonomous for the routine, cautious for the important.

Memory: Honcho as the hippocampus

Memory is enabled with user_profile_enabled: true and the Honcho provider. The agent writes every observation into the key structure described earlier, and reads from it as its primary context source. The Rust → FRB → Flutter bridge post explained how the on-device AI engine connects; at the agent layer, the same split applies — the deterministic analysis runs close to the phone, while the judgment and narrative layers run inside the agent's memory-rich session.

How the Brain Reaches the Phone: Bridges and Feedback Loops

The agent runs where the owner runs it — on their own machine, reachable only through authenticated channels. Two bridges connect it to the app.

The Python bridge: profile and gateway management

hermes-bridge.py is an HTTP service (default port 8742) that the Flutter app calls to install or update the profile, start and stop the gateway, and query status. Its endpoints are narrow: /health, /profile/list, /profile/status, /profile/install, /gateway/start, /gateway/stop, /gateway/status, and /worker/status (which probes connectivity to the Cloudflare Worker). Every request requires a Bearer token compared with hmac.compare_digest — constant-time comparison, the same discipline the care codes use. On Linux the bridge runs as a systemd service (puppal-bridge.service), which is how the phone can manage an agent that is technically not running on the phone at all.

The Rust bridge: skill visibility in the UI

On the app side, the skillListProvider in the Flutter code fetches PiSkillInfo entries through the FRB bridge by calling listSkills(dataDir) — the same distribution directory the agent manages. The app uses this list for the skill tree page, the badge in the top bar, and the check-in diff: after a check-in, Puppal compares the skill set before and after via diffNewSkills and shows the owner which new skills their pet's care has unlocked. The photo analysis result is parsed into a typed PhotoAnalysisResult model with v2Posture, v4Expression, v5Environment, c7Anomalies, anomalyScore, needsOwnerAttention, and a message the owner can edit — with isEdited/editedMessage tracking so the agent's version and the human's version are never confused. That model aligns field-for-field with the JSON contract the photo-analyze skill promises, which is how a Python-written agent and a Dart-written app can agree on what a check-in means.

Conclusion: Why Hermes Agent Skills Scale into a Care Network

Step back from the components and the architecture is surprisingly simple: PupPal's Hermes Agent skills divide the product into five single-purpose competencies, ship them as a versioned profile distribution, trigger them through webhooks and cron, and connect them through one shared Memory contract. That is the whole brain — and it is exactly why the roadmap keeps expanding without a rewrite. Foster proxy check-in was added as the care-monitor skill plus a routing rule. The family-circle feature under development will be a new skill plus new webhooks, reusing the same profile, Memory, and bridge plumbing. Adoption, the next milestone after check-in and foster care, will need a profile handover flow more than a new engine.

For the owner, all of this machinery collapses into one reassuring fact: whether it is their own photo at breakfast, a caregiver's photo at noon, the voice at nine, or the review at nine at night, the same well-defined brain is watching, recording, and — only when it really matters — speaking up. That is what Hermes Agent skills deliver for PupPal: an AI pet check-in companion that is powerful enough to think, honest enough to say when it cannot, and structured enough to grow.