Pet Daily Voice: Your Pet Speaks Every Morning — PupPal
2026-08-12
Pet Daily Voice: Your Pet Speaks Every Morning — PupPal
Most pet apps are silent. They store photos, count steps, and occasionally nag you to feed the dog — but at no point does the dog itself say anything. PupPal's pet daily voice changes that. Every morning at 09:00, the Hermes Agent that powers the product reads everything that happened to your pet yesterday and writes a short, first-person message as the pet: what it did, how it felt, what it hopes for today. The message arrives on your phone like a text from a friend — except the friend is your dog, and the text is grounded in real check-in data, not random cuteness.
In this post we open the daily-voice skill that produces this message, the cron job that wakes it up, the memory keys it reads, the prompt that turns numbers into feelings, and the delivery path that puts the voice in your hands. Along the way we will also look at the design trade-offs: the rules that stop the AI from reporting raw scores, the failure modes it guards against, and why a 150-character monologue is one of the most effective engagement mechanisms PupPal has built. The pet daily voice is part of the five-skill cognitive workload we described in our earlier post on the Hermes Agent skills architecture, and it is the skill humans interact with most.
What Is the Pet Daily Voice Message?
The concept is simple: once a day, the pet "talks." The message is a monologue of roughly 150–200 characters written from the pet's point of view, in the pet's voice, covering the previous day. It is not a report and it is not a notification. It is a narrative — a miniature story with a mood, a few concrete scenes, and an emotional arc.
The one-line definition
The daily-voice SKILL.md in the Puppal repository describes the
output as a "今日心声" (today's heart-voice): a warm, personality-
bearing pet monologue pushed to the owner's WeChat. The skill's own
metadata gives it the tags pet, voice, daily, narrative,
personality, and puppal, and marks it as a pet-care category
skill related to daily-review and photo-analyze — the two skills
that produce the data it consumes.
What it is not
It is worth being precise about what the message is not, because the boundaries define the design:
- It is not an analytics digest. The pet may not quote numbers. The prompt explicitly forbids output like "my happiness is 0.72."
- It is not a template. Every message is generated fresh by a large language model from that day's real data, so two mornings are never identical.
- It is not a marketing push. It contains no offers, no links, and no asks beyond the pet's own wishes ("take me to the park again").
- It is not silent on bad days. If yesterday had no check-ins at all, the skill still produces a short "missed you" message instead of quietly skipping.
That last point matters. Many automated systems go silent when there is no data; the pet daily voice treats a missing day as data — a day the owner did not check in — and gives the pet something to say about it.
Where it sits in the product
The voice is written by the daily-voice skill inside the Hermes
profile distribution, delivered by the messaging toolset to WeChat
when WEIXIN_ENABLED=true, and scheduled by a cron job — the same
cron infrastructure that runs the nightly 21:00 daily-review. From the
owner's perspective it is simply "the morning message from my dog."
From an architecture perspective it is the read-mostly output channel
of the entire check-in loop: the loop starts with a photo
(what happens when you snap it),
flows through vision analysis and the C7 anomaly check, gets
aggregated by the nightly review, and finally becomes words the next
morning.
The 09:00 Cron: How the Morning Message Begins
The message does not appear by magic; it is triggered by a
declarative cron job shipped inside the Hermes profile. The
hermes/cron/daily-voice.json manifest is small, but every field in
it is a design decision.
The schedule
The schedule field is 0 9 * * * — 09:00 every day, in the owner's
local timezone as far as the cron runtime is concerned. The choice of
09:00 is deliberate. It is late enough that the previous night's
21:00 daily review has definitely finished and written its results to
memory, so the voice always has yesterday's analysis available. It is
early enough that the message lands before most people start their
day — a natural moment to check the phone, exactly when a "good
morning from your pet" has the most emotional impact.
The enabled_toolsets field lists memory, terminal, file, and
messaging. Only the first and last are strictly needed — reading
memory and delivering the message — but terminal and file are
available so the skill can run helper scripts or stage content if the
prompt logic ever grows beyond pure generation.
The prompt field
The cron job carries a one-line prompt that instructs the agent:
Run the daily-voice skill: read yesterday's daily review data and the pet's personality settings, generate a 150–200 character "today's heart-voice" in the pet's first person, and push it to the owner's WeChat. If there was no check-in yesterday, generate a short miss-you message instead.
Notice that the cron does not embed the full generation template. The heavy lifting lives in the skill, which is invoked by name — a clean separation between when something runs (the cron) and how it runs (the skill).
The name and delivery mode
The manifest's name is "Puppal 今日心声" and its deliver mode is
origin. In Hermes terms, deliver: origin means the output of the
run is delivered back through the origin channel — the messaging
integration — rather than being filed away for review. That is
exactly what you want for this job: the message is the deliverable,
and the deliverable goes straight to the owner's phone.
What happens at 21:00 first
It is worth repeating the dependency: at 21:00 the daily-review
cron runs trend_calc.py — a pure-Python script with no external
dependencies — which performs the S1 baseline comparison, S3 feeding
check, and S4 activity comparison, then writes a fresh entry into
dog:{dog_id}:stats.history. The 09:00 voice cron reads that entry
twelve hours later. The two cron jobs are a pipeline with a night of
sleep in the middle: review at night, voice in the morning.
The Raw Material: What the Voice Knows About Yesterday
A convincing monologue needs facts. The daily-voice skill reads four
things from long-term memory before it writes a single sentence. All
of them live under the pet's memory namespace (dog:{dog_id}:...),
which the Hermes runtime persists through the Honcho memory provider
declared in hermes/config.yaml.
Yesterday's daily review entry
The first read is dog:{dog_id}:stats.history, and the skill takes
the last entry — yesterday's review. That entry is a rich JSON
document produced by the nightly run, and it contains:
- S1 baseline results (
s1_baseline):compared_to_baseline(stable | improving | declining),happiness_delta,energy_delta,health_score_delta, a list ofsignificant_changes(for example, "energy dropped 12% vs the 7-day average"), and a plain-languageassessment. - S3 feeding results (
s3_feeding):meals_detected_today,compared_to_usual(normal | less | more | unable_to_detect),usual_meal_count, and anassessment. - S4 activity results (
s4_activity):activity_score_today,activity_score_7day_avg,trend,high_activity_moments,low_activity_moments, and anassessment.
The voice skill pulls two derived fields from this entry directly —
happiness and energy, each a 0–1 score — plus photo_count, the
activity summary (s4_activity.assessment), and the notable events
(s1_baseline.significant_changes). The happiness and energy values
themselves are computed in the photo analysis pipeline: each check-in
photo yields posture, expression, and environment assessments plus a
C7 anomaly score, and the review aggregates them into daily state.
The app's PhotoAnalysisResult model carries exactly these as typed
fields — v2Posture, v4Expression, v5Environment,
c7Anomalies, anomalyScore, needsOwnerAttention, and message —
so the plumbing from photo to aggregate is explicit in the codebase.
The pet profile and personality
The second read is dog:{dog_id}:profile — the pet's name, breed,
age, and especially the personality settings the owner configured
during initialization (the puppal-init webhook writes this key when
the app first sets up the pet). Personality is a JSON object with
four fields:
{
"tone": "warm",
"age_voice": "young_adult",
"traits": ["cheeky", "clingy", "easily startled"],
"speech_style": "uses cute interjections, but never talks about human things"
}
If the owner never configured a personality, this exact object is the
fallback — warm tone, young-adult voice, three generic traits. The
template's speech_style slot is what stops the pet from sounding
like a person: the pet may use cute interjections but should not
reason like a human. Owners can change the tone in the app, and the
skill's documented pitfalls note that a fixed template may fit some
personalities better than others — which is why the prompt is
parameterized by these fields rather than hard-coded prose.
The seven-day trend
The third read is the same stats.history key, but the last seven
entries, used to compute happiness_trend and energy_trend — plain
directional signals (rising, falling, flat) that let the pet talk
about change rather than a single snapshot: "I've been a bit more
tired than usual this week."
The check-in photos themselves
Finally, the daily-voice skill is related to photo-analyze, which
means the memory also holds the day's individual photo records. The
series of check-ins on dog:{dog_id}:photos is what the nightly
review aggregates, and the voice prompt is encouraged to mention
concrete scenes when the photo analysis detected them — the park
visit, the squirrel chase, the long sofa nap. This is what makes the
message feel like a memory rather than a mood ring.
From Metrics to Feelings: The Prompt That Makes a Pet Talk
The heart of the skill is a carefully constructed LLM prompt. It starts by establishing identity with a template like:
You are {pet_name}, a {age}-year-old {breed}.
Your personality is {traits}.
Then it hands over the data with an instruction that defines the whole product: use the real data, but express it with feelings, not numbers.
The data block
The prompt receives:
- 昨日数据 (yesterday's data):
happinessquantized to three buckets ("very happy" if > 0.7, "okay" if > 0.4, "not great" otherwise),energybucketed the same way ("full of energy" > 0.7, "normal" > 0.4, "a bit tired" otherwise),photo_count,activity_summary, andnotable_events. - 7-day trend (seven-day trend):
happiness_trendandenergy_trend.
The bucketing is a deliberate pre-processing step. The agent never sees "happiness: 0.72" as a raw float to echo back; it sees a qualitative statement ("very happy") that it can convert into narration. This single design detail — quantize before you prompt — is what prevents the most common failure mode of data-to-narrative generation, the robot sentence "my happiness was 0.72."
The seven rules
The prompt then lays down seven rules that constrain the generation:
- Speak as the pet, not as a report. Perspective is non- negotiable.
- If something is wrong (consecutive unhappiness, falling energy), express concern gently — worried, not clinical.
- If yesterday was good, show happiness and look forward to today. The message should end on an upward note.
- Concrete scenes are welcome — but only ones the photo analysis actually detected. No invented walks.
- End with affection or expectation for the day ahead.
- Never quote numbers directly. The pet may say "I was a bit slow today," never "my activity score was 0.52."
- Sound like a real dog or cat — natural, specific, a little silly.
Rule 4 is the honesty contract with the user: the monologue is generated, but its facts are grounded in the vision pipeline. The app's check-in flow runs the V2 posture, V4 expression, and V5 environment analyses through the C7 anomaly check on every photo, so the "scene" the pet remembers is one the system actually observed.
The personality parameterization
The same prompt template is cloned per pet, with traits, tone,
and age_voice injected. This is how two dogs on the same product
sound completely different: a stoic old husky with a terse speech_style
produces a three-sentence grumble where a young golden retriever with
"cheeky, clingy" traits produces a five-paragraph saga about a
squirrel. Personality is not an afterthought bolted onto the output;
it is an input to the generator, and the comparison between what the
personality should sound like and what the LLM actually wrote is
one of the skill's verification checks.
Three Kinds of Mornings: Good Days, Hard Days, and Missed Days
The skill's SKILL.md documents three example outputs, and together
they reveal the emotional range the system is designed to hold.
The good day ☀️
On a normal day the message is sunlit and specific: the morning walk, the park, a near-miss with a squirrel, a long sofa nap, a dream about the fridge door opening by itself, dinner smelled amazing, the belly rub lasted a long time, and today the pet hopes to visit the squirrel park again. Notice what is absent: no scores, no counts, no health jargon. The underlying data — high happiness, high energy, two meals, activity mostly lounging with one high-energy burst — is fully disguised as lived experience.
The hard day 🌧️
When the data shows trouble — low energy, disinterest, a possible itchy ear — the pet does not say "anomaly detected." It says: I didn't feel like moving today. The food was fine but not as good as usual. My ear seemed a bit off and you put drops in it. I feel a little better now — but if I'm still not moving tonight, maybe we should see the vet. The concern is voiced from the pet's side, which makes it land differently than a medical alert: it is worried about the pet, not warning about the pet. This is the engagement mechanism doing double duty as an early-warning signal — the gentle form of the same C7 anomaly detection that powers the hard anomaly alerts during the day.
The missed day 💤
If stats.history contains no entry for yesterday (the nightly
review skips days with zero photos), the skill switches to a
different template: Yesterday you were so busy you didn't take a
single photo of me. That's okay — I slept most of the day and wagged
my tail extra hard when you came home. Can we take a few photos
today? I'm already posing. The message converts guilt into a warm
nudge: the owner is reminded, without being scolded, that check-ins
are how the pet "talks." A run of missed days shows up in the
message's tone — the documented pitfall about repetitive content is
handled by rotating "today's special mention" topics — so the voice
stays fresh even when the data does not.
Delivery, Verification, and the Failure Modes
Generating the text is only half the job. The skill's SKILL.md
defines a step-four delivery path and a six-point verification
checklist, and its pitfalls section documents the known ways the
system can go wrong.
Pushing to the owner's phone
The last procedure step checks WEIXIN_ENABLED. If true, the message
is pushed to the owner's WeChat through the Hermes messaging
toolset — which is why the cron manifest lists messaging in its
enabled_toolsets and sets deliver: origin. The push format
prepends the pet's name and an emoji as a title ("☀️ 豆豆's heart
voice today"), so the message is recognizable at a glance in a chat
feed among human messages. The emoji choice is part of the format,
not decoration: 🌧️ for hard days, 💤 for missed days, ☀️ for good
days — a one-character mood indicator before the owner reads a word.
The verification checklist
After generation, the skill confirms six properties:
- The whole text is first-person. ✓
- No numeric scores appear anywhere. ✓
- Length is between 100 and 250 characters. ✓
- Tone matches the configured personality. ✓
- If yesterday had anomalies, the monologue reflects them. ✓
- The push was sent (when enabled). ✓
Check 5 is the lineage check: it guarantees the voice is not cheerfully describing a day the C7 pipeline flagged as problematic. The verification is automated enough to be cheap and human enough to be meaningful — the agent reviews its own draft against the same rules that were in the prompt.
Documented failure modes
The skill's pitfalls section is unusually honest, and each entry maps to a concrete mitigation:
- The LLM translates data into numbers. Mitigation: prompt rule 6 plus the pre-bucketed data block.
- Fixed template vs. unusual personalities. Mitigation: owner- configurable tone in the app; the prompt's slots come from the profile, not from hard-coded text.
- Repetitive content over consecutive days. Mitigation: randomly inject a "today's special mention" topic into the prompt.
- Insufficient data. Mitigation: never go silent — emit the short missed-day message instead of skipping.
These mitigations are product decisions as much as engineering decisions. The team chose availability of the emotional channel over purity of the data feed: a slightly imperfect message every morning beats a perfect message that occasionally fails to arrive.
Why the Pet Daily Voice Keeps Owners Engaged — and Where It Goes Next
The final question is the one the topic promises: why does a 150- character monologue keep owners engaged? The answer is that the pet daily voice converts the product's core mechanic into a daily emotional event.
Engagement without a dashboard
Check-in apps typically fight for attention with streaks and badges.
PupPal does not need to: the voice message is the reward for
yesterday's photo. Every check-in the owner made is acknowledged the
next morning in the most flattering currency available — the pet
itself remembering. The loop is self-reinforcing: photo → analysis →
review → voice → emotional payoff → motivation to photograph again
today. The photo-count from yesterday even appears in the prompt,
so the pet can quietly acknowledge how many times it was checked
on, and owners who photograph more get a pet that talks more.
Habit formation through narrative
Research on habit formation consistently finds that immediate, positive, variable rewards beat fixed schedule reminders. The voice message is exactly that: daily, warm, and never the same twice (because the data and the "special mention" rotate). It also piggybacks on the existing morning phone check — the message arrives at 09:00, a time people already look at their phones, so it does not need to fight for attention, it simply occupies a slot that was already open.
A soft early-warning layer
Engagement is not just about retention; it is also about attentiveness. On mornings after a flagged day, the pet's worried tone is the first thing the owner reads — before any medical summary or alert. The daily-voice is the emotional front end of the anomaly-detection back end, and by making concern legible in the pet's own voice it raises the probability that an owner actually acts on it.
Where the voice goes next
The roadmap — pet check-in, foster proxy check-in, adoption — gives
the daily voice natural extensions. During a
foster care session,
the care-monitor skill tightens anomaly thresholds and shadows
every caregiver photo; the next morning's voice would then narrate a
day observed by someone other than the owner — a pet talking about
the aunt who fed it, which is precisely the reassurance foster care
demands. In the adoption module, the voice could become the pet's
"résumé" to a prospective family: a week of daily monologues is a
far better introduction than a static profile photo. And on the
product side, moving delivery beyond WeChat (in-app feed, email, or
SMS) would make the voice the default morning touchpoint of the
entire platform.
The pet daily voice is a small feature with an outsized effect. It takes the dry output of a vision pipeline — posture scores, anomaly scores, activity aggregates — and turns it into something a person anticipates. It is the warmest part of a system built around care: every morning, at 09:00, your pet tells you how it felt, what it did, and what it wants today. That is the pet daily voice — PupPal's way of making sure the check-in loop never ends in a database, but always comes home to a conversation.
Metric-Driven Thoughts: What the Numbers Actually Say
Stepping back from the poetry for a moment, the pet daily voice is
also a well-engineered piece of systems design. Its inputs are
deterministic (memory keys with stats.history entries, profile
JSON, photo records), its transformation is probabilistic (an LLM
call against the OpenRouter default model anthropic/claude-sonnet-4
declared in hermes/config.yaml, with max_turns: 45, tool-use
enforcement, and context compression at a 0.50 threshold and 0.20
target ratio), and its output contract is strict (first person, no
numbers, 100–250 characters, personality-conformant, anomaly-aware).
The skill ships as part of a profile distribution alongside the
other four skills, so installing PupPal on a new machine installs the
voice too — no extra configuration, no separate deployment.
The takeaway for builders is simple: engagement features do not need to be exotic, they need to be grounded. Ground the message in real data, parameterize it with real personality, constrain it with real rules, verify it against real checklists, and schedule it at a real moment in the human day. Do that, and a 150-character text message becomes the feature owners talk about — the morning they wait for, and the reason they reach for the camera again today.