Hermes Agent Skills: Inside the AI Brain of PupPal — PupPal
2026-08-11
Hermes Agent Skills: Inside the AI Brain of PupPal — PupPal
If you have been following this blog, you already know the PupPal
mantra: photo = check-in, share = care. A pet owner snaps a
photo, and a few seconds later the app shows posture, expression,
environment, and an anomaly score. A caregiver types a care code and
PIN, and suddenly they can check in on a pet they have never met. The
posts so far have described these flows from the outside — what
happens when you press the shutter, what the C7 anomaly engine
checks, how care codes stay secure. This post opens the box and looks
at the part that actually thinks: the Hermes Agent skills
architecture inside PupPal. Five skills — photo-analyze,
daily-review, daily-voice, care-monitor, and care-handbook —
form the complete cognitive workload of the product, and they are
shipped to the user's own machine as a Hermes profile distribution.
The word "brain" is not a metaphor here. The agent is a genuinely autonomous software entity: it watches webhooks, wakes up on a cron schedule, keeps a long-term memory, and decides — within strict rules — what to analyze, what to write, and what to push to the owner's phone. Understanding how that brain is organized explains almost every behavior the earlier posts documented. It also explains why PupPal can extend to foster care, family circles, and eventually adoption without redesigning the product: because each new capability is just a new skill in a distribution that already knows how to install, trigger, and govern itself.
Why Puppal Put an Agent, Not a Pipeline, in Charge
Before diving into the five skills, it is worth asking why PupPal uses an agent at all. A conventional implementation would be a chain of API calls: upload photo, call a vision model, run a rule engine, update a database, send a push notification. That pipeline works for a single deterministic flow. It fails the moment you add scheduling, personality, context, and judgment.
From stateless calls to a stateful caretaker
A pipeline has no memory between calls. PupPal needs memory everywhere:
the anomaly check compares today's photo against the last seven days;
the nightly review needs the whole day's photos plus the previous
month of statistics; the morning voice needs yesterday's review to
speak as the pet. An agent, by contrast, lives inside a persistent
session with tools, memory, and a schedule. The Hermes runtime that
PupPal relies on provides exactly that substrate, and the profile
distribution declares it explicitly. The config.yaml in the
hermes/ directory of the Puppal repo sets max_turns: 45, enables
tool-use enforcement, turns on context compression at a 0.50
threshold with a 0.20 target ratio, and — critically — enables the
memory toolset with the Honcho provider for long-term storage. The
agent is not a script that runs once; it is a resident caretaker
that can be asked to do work at any moment, with the full history of
the pet available in its working context.
What "photo = check-in, share = care" means for the brain
The two halves of the mantra map onto two halves of the skill architecture. "Photo = check-in" is the observation loop: a photo arrives, the agent looks, judges, records. "Share = care" is the coordination loop: the owner initiates a care session, the agent writes a handbook, then watches every caregiver photo with stricter eyes. Both loops share the same infrastructure — the same Memory keys, the same webhook channel, the same vision toolset — but they are implemented as separate skills so that each can be tuned independently. That separation is the architectural core of the product: a pet care AI is only as good as the boundaries it draws between its responsibilities.
The profile distribution: the brain as a shippable unit
The most distinctive decision in the architecture is that the brain
is distributed as a Hermes profile distribution, not embedded in
the app. The distribution.yaml manifest names the profile
(puppal, version 1.0.0), declares the Hermes runtime requirement
(hermes_requires: ">=0.13.0"), lists the environment variables it
needs (OPENAI_API_KEY, WORKER_WEBHOOK_SECRET,
PUPPAL_WORKER_URL, and optional WEIXIN_ENABLED), and claims
ownership of exactly the files it manages: SOUL.md, config.yaml,
mcp.json, skills/, and cron/. Installing the profile is a
single command — hermes profile install ./hermes — and the same
manifest is what the bridge service uses to update the profile in
place. The app does not need to embed prompts or model calls; it
talks to an agent that can be versioned, upgraded, and even replaced
independently of the Flutter code.
The Five Skills and What Each One Owns
The entire cognitive workload of PupPal is divided among five
skills. Each skill is a directory containing a SKILL.md — the
markdown file that is both documentation and executable instruction
for the agent — plus optional Python scripts for the deterministic
parts. The distribution's distribution_owned list means Hermes
treats these directories as managed state: installs and updates
overwrite them atomically.
| Skill | Trigger | Responsibility |
|---|---|---|
| photo-analyze | Webhook puppal-photo (source=owner) | Vision V2+V4+V5, C7 anomaly check, state update |
| daily-review | Cron 21:00 | S1 baseline, S3 feeding, S4 activity, health score, trend alerts |
| daily-voice | Cron 09:00 | First-person pet monologue from yesterday's real data |
| care-handbook | Webhook puppal-care-create | Structured care instructions for a stranger |
| care-monitor | Webhook puppal-photo (source=caregiver) | Stricter analysis during care sessions, instant alerts |
photo-analyze — the eye
photo-analyze is the busiest skill and the one documented most
thoroughly in the photo check-in deep
dive.
Its procedure is worth summarizing here because it shows the
skill-template pattern: verify the request, gather context, call the
vision toolset, run the deterministic script, update Memory, return a
structured result. One vision call produces V2 posture, V4
expression, and V5 environment; the local anomaly_check.py script
then computes C7 anomalies against the last seven days of history.
The skill's own front matter declares related_skills: [daily-review, care-monitor], which is how the agent knows where to
route follow-up questions.
The skill also documents a deliberately cheaper App Checkin
mode for on-device check-ins: images are compressed to a 1568px
long edge, files over 4MB are forced to JPEG, and a single merged
prompt returns only contract JSON. Text check-ins skip vision
entirely and go through classifyActivity. The threshold logic is
spelled out in the skill itself: anomaly_score > 0.5 flags the
message with a warning, > 0.7 triggers a standalone push, a
critical anomaly pushes immediately, and a single-day health drop
of more than 0.15 is highlighted. Every number in this product
lives in a skill file where the agent can read it — and where
reviewers can audit it.
daily-review — the memory of trends
At 21:00 every night, the daily-review cron job fires and the
skill aggregates the day. It reads today's photos from
dog:{dog_id}:photos, the current state from dog:{dog_id}:state,
and the statistical history from dog:{dog_id}:stats, then runs
trend_calc.py for the three nightly analyses: S1 baseline
comparison, S3 feeding confirmation, and S4 activity comparison.
The health score is a weighted formula — happiness 0.25, inverted
anomaly average 0.30, energy 0.15, feeding normality 0.15, activity
versus baseline 0.15 — and the skill defines the alert ladder: three
consecutive days of health decline exceeding 0.2 is a high-severity
push, three days of energy under 0.3 is high, a single bad-happiness
day is medium, and an anomaly_score above 0.5 is medium. When
everything is normal, the report is stored silently; only real
signals reach the owner's phone, which is exactly the notification
discipline the product promises.
daily-voice — the voice
At 09:00, daily-voice turns yesterday's numbers into a pet's
first-person monologue. The skill is meticulous about its rules: no
raw numbers ("my happiness is 0.72" is forbidden), the personality
comes from the pet profile's tone, age_voice, and traits, and
the length is held to roughly 150–200 characters of Chinese text.
The prompt is constructed from real data — yesterday's review, the
seven-day trend, notable events from significant_changes — but the
output is feelings, not statistics. If yesterday had zero check-ins,
the skill does not go silent; it produces a short, wistful message
about the owner being too busy to take photos. That behavior is
deliberate: engagement is maintained by the agent showing up
every single day, even when there is nothing to report.
care-handbook — the briefing
When the owner starts a care session, care-handbook turns the pet
profile into instructions a stranger can follow. The fascinating
part is what the LLM is not allowed to do. Feeding amounts,
medication names and dosages, vet contact information, and emergency
contacts are copied verbatim from Memory — the skill explicitly
instructs the agent to verify field-by-field equality and never let
the model rewrite a number. The LLM is only permitted to generate
the behavior suggestions, the check-in instructions, and a closing
encouragement, and the whole handbook is capped at 2000 characters.
Emergency contact details are never shown to the caregiver at all;
the handbook says to contact the vet instead. This is a template for
how to use LLMs responsibly around safety-critical facts: let the
model write prose, never let it write truth.
care-monitor — the vigilant stand-in
care-monitor is photo-analyze with the safety margins cut in
half. When source=caregiver, the webhook routes to this skill
instead: all thresholds drop by 40% (push threshold 0.7 → 0.42,
health-drop alert 0.15 → 0.09, low-mood streak 5 photos → 3), and
alerts go to the owner immediately instead of waiting for the
nightly review. It adds care-specific checks: the first anomaly of
the session is flagged on sight, a repeated anomaly type escalates
severity, caregiver notes containing keywords like "not eating" or
"vomiting" trigger an immediate high alert, and an environment that
moves from home to elsewhere is flagged as adaptation stress. The
foster care monitoring post
documents this in depth; the skill file itself states the design
philosophy bluntly: it would rather alarm once too often than miss
once. Every caregiver photo is recorded into the care session record
at dog:{dog_id}:care_sessions.{care_session_id}.caregiver_photos
with its anomaly_score and whether the owner was alerted.
How the Skills Are Triggered: Webhooks, Cron, and Routing
Skills never run on their own initiative. Everything is driven by two trigger mechanisms, and the routing between them is explicit in the webhook definitions.
Webhook-driven entry points
The Cloudflare Worker is the only caller of the agent's webhook
endpoints, and every call carries the X-Webhook-Secret header
checked against WORKER_WEBHOOK_SECRET — an unauthorized call
returns 401 before any analysis happens. The four webhooks that
matter for the skill brain are:
puppal-init— creates the pet profile and initial state in Memory when the owner first sets up the app.puppal-photo— carriesphoto_url,timestamp,source, an optionalcaregiver_note, andcare_session_id. Thesourcefield is the router:owner→photo-analyze,caregiver→care-monitor.puppal-care-create— opens a care session and invokescare-handbook, returning the handbook JSON to the Worker.puppal-care-end— closes the session with areasonofmanualorexpired, sets the session status toended, and invalidates the care code immediately.
The webhook-driven design is what keeps the agent decoupled from the app: the Worker never needs to know which model or which prompt is in use, and the agent never needs to know how the app renders results.
Cron-driven jobs
Two cron jobs live in the profile's cron/ directory as JSON files
with a schedule, a prompt, the skills to load, and an
enabled_toolsets allowlist. daily-review.json runs at
0 21 * * * and loads memory, terminal, file, messaging, vision,
and web tools; daily-voice.json runs at 0 9 * * * with a
smaller toolset. Notice what is missing from the voice cron:
terminal. The voice skill has no Python scripts, so it does not
need one. The allowlist is a quiet but real security feature —
each scheduled job gets exactly the tools its skill needs and
nothing more.
Shared state: the Memory contract
All five skills communicate through the same Memory keys, which makes the brain coherent without any inter-skill messaging. The contract is simple and stable:
dog:{dog_id}:profile— name, breed, age, weight, feeding, medications, allergies, behavior, vet, emergency contact.dog:{dog_id}:state— rolling happiness and energy (updated with an exponential moving average that weights the newest reading at 0.3),health_score,last_photo_at,photo_count_today.dog:{dog_id}:photos— append-only photo history with every analysis result attached.dog:{dog_id}:statsandstats.history— daily aggregates used by the review, the voice, and the trend checks.dog:{dog_id}:care_sessions— active and ended sessions, each with its handbook timestamp and caregiver photo log.
Because the state is versioned and append-only, the agent can answer questions like "is energy declining?" by reading the same data the nightly review writes — no pipeline bookkeeping required.
Inside a Skill: SKILL.md as the Unit of Behavior
A skill is not a plugin binary or a compiled module. It is a markdown file with YAML front matter, plus scripts. That choice looks unusual until you realize what it buys: the agent itself can read its own operating instructions, and humans can review them in a pull request. The C7 anomaly detection post described the engine; here is the container it lives in.
The front matter contract
Every SKILL.md opens with name, description, platforms,
version, author, license, and a metadata.hermes block that
carries tags, category, command prerequisites (python3), and
related_skills. The description is written to be read by the
agent's skill tooling — it is a functional summary of when to use
the skill, not marketing copy. The required_environment_variables
block tells the runtime which secrets must exist before the skill
can be trusted to run. This metadata is what lets Hermes expose the
skills to the app: the Flutter side lists them through the Rust
bridge's listSkills call, which reads the same distribution
directory and returns PiSkillInfo entries.
The body structure: procedure, thresholds, pitfalls
Each skill body follows the same disciplined outline: a data-flow
diagram, "When to Use" (explicitly stating the skill is triggered
only by webhook or cron, never called directly by the user),
numbered procedure steps, a thresholds table, a pitfalls list, a
verification checklist, and a quick-reference table. The pitfalls
sections are unusually honest — they enumerate the failure modes the
designers actually expect: low-confidence vision on blurry photos,
indoor/outdoor misclassification in low light, no baseline for new
pets in their first seven days, vision timeouts on large photos, and
Honcho memory connectivity issues. The verification checklists are
post-conditions the agent must confirm before it reports success,
such as "Vision returned valid JSON", "the C7 script exited 0", and
"last_photo_at was updated". A skill file is a contract with the
agent about what done looks like.
Deterministic cores in a stochastic world
The vision and writing parts of the pipeline are LLM work, but the
parts that must be exact are Python. anomaly_check.py implements
C7 as a pure rule engine over the vision output and history; it
takes a --mode care flag that activates the stricter care checks.
trend_calc.py computes baselines, feeding counts, and activity
comparisons from JSON on stdin and returns JSON on stdout. Keeping
these deterministic keeps the system auditable: the same photo and
the same history always produce the same anomaly score, regardless
of which model version happens to be serving the vision call.
What the Agent Brain Knows — and Refuses to Fake
A brain is more than its reflexes. The Puppal agent's judgment is shaped by three layers: the SOUL file, the configuration, and the memory system.
SOUL.md: personality as policy
SOUL.md defines who the agent is: an observer, a recorder, an
alarm, a narrator, and a care coordinator — the five roles map
one-to-one onto the five skills. Its core principles are worth
quoting in spirit: data honesty is the first principle ("never
fabricate analysis results"); output is threshold-driven (normal
check-ins get one-line confirmations, anomalies get detail and
advice); trends matter more than single points ("one photo's
anomaly may be a misjudgment; three days of trend is the real
signal"); care periods are more sensitive by design; and the
owner's notification experience is protected by merging routine
output into the nightly review. These principles are not decorative
— the skills encode them into thresholds and push rules, so the
personality is enforced mechanically, not merely suggested.
Config: the knobs that govern judgment
config.yaml reveals the operational envelope: the default model is
anthropic/claude-sonnet-4 via OpenRouter with a 200,000-token
context, delegation subtasks run on the same model with up to 30
iterations, tirith_enabled turns on the security layer, and
approvals run in smart mode — meaning the agent can act on routine,
well-specified work without pestering the owner, while still
stopping for confirmation on genuinely consequential actions. The
compression settings (threshold 0.50, target ratio 0.20) keep long
conversations within context by summarizing older turns before they
overflow. Together these knobs define how autonomous the brain is
allowed to be — and the answer is: autonomous for the routine,
cautious for the important.
Memory: Honcho as the hippocampus
Memory is enabled with user_profile_enabled: true and the Honcho
provider. The agent writes every observation into the key structure
described earlier, and reads from it as its primary context source.
The Rust → FRB → Flutter bridge
post
explained how the on-device AI engine connects; at the agent layer,
the same split applies — the deterministic analysis runs close to
the phone, while the judgment and narrative layers run inside the
agent's memory-rich session.
How the Brain Reaches the Phone: Bridges and Feedback Loops
The agent runs where the owner runs it — on their own machine, reachable only through authenticated channels. Two bridges connect it to the app.
The Python bridge: profile and gateway management
hermes-bridge.py is an HTTP service (default port 8742) that the
Flutter app calls to install or update the profile, start and stop
the gateway, and query status. Its endpoints are narrow: /health,
/profile/list, /profile/status, /profile/install,
/gateway/start, /gateway/stop, /gateway/status, and
/worker/status (which probes connectivity to the Cloudflare
Worker). Every request requires a Bearer token compared with
hmac.compare_digest — constant-time comparison, the same
discipline the care codes use. On Linux the bridge runs as a
systemd service (puppal-bridge.service), which is how the phone
can manage an agent that is technically not running on the phone at
all.
The Rust bridge: skill visibility in the UI
On the app side, the skillListProvider in the Flutter code fetches
PiSkillInfo entries through the FRB bridge by calling
listSkills(dataDir) — the same distribution directory the agent
manages. The app uses this list for the skill tree page, the badge
in the top bar, and the check-in diff: after a check-in, Puppal
compares the skill set before and after via diffNewSkills and
shows the owner which new skills their pet's care has unlocked. The
photo analysis result is parsed into a typed PhotoAnalysisResult
model with v2Posture, v4Expression, v5Environment,
c7Anomalies, anomalyScore, needsOwnerAttention, and a message
the owner can edit — with isEdited/editedMessage tracking so the
agent's version and the human's version are never confused. That
model aligns field-for-field with the JSON contract the
photo-analyze skill promises, which is how a Python-written agent
and a Dart-written app can agree on what a check-in means.
Conclusion: Why Hermes Agent Skills Scale into a Care Network
Step back from the components and the architecture is surprisingly
simple: PupPal's Hermes Agent skills divide the product into five
single-purpose competencies, ship them as a versioned profile
distribution, trigger them through webhooks and cron, and connect
them through one shared Memory contract. That is the whole brain —
and it is exactly why the roadmap keeps expanding without a rewrite.
Foster proxy check-in was added as the care-monitor skill plus a
routing rule. The family-circle feature under development will be a
new skill plus new webhooks, reusing the same profile, Memory, and
bridge plumbing. Adoption, the next milestone after check-in and
foster care, will need a profile handover flow more than a new
engine.
For the owner, all of this machinery collapses into one reassuring fact: whether it is their own photo at breakfast, a caregiver's photo at noon, the voice at nine, or the review at nine at night, the same well-defined brain is watching, recording, and — only when it really matters — speaking up. That is what Hermes Agent skills deliver for PupPal: an AI pet check-in companion that is powerful enough to think, honest enough to say when it cannot, and structured enough to grow.