Pet Daily Health Review: The 21:00 Ritual That Catches Health Drift — PupPal

2026-08-16

Pet Daily Health Review: The 21:00 Ritual That Catches Health Drift — PupPal

A single photo tells you how your dog looked at 14:32 on a Tuesday. It cannot tell you that this was the fourth quiet day in a row, that energy has been sliding for a week, or that the last three nights showed no eating posture at all. That is the difference between a snapshot and a trend — and it is exactly the gap the pet daily health review in PupPal exists to close. Every night at 21:00, the Hermes Agent that powers the product reads all of the day's photo check-ins, compares them against the previous seven days, computes a health score, and decides whether anything is drifting quietly toward trouble.

This post opens the nightly review: the cron job that wakes it, the three analyses it runs (baseline comparison, feeding confirmation, activity comparison), the weighted health score it produces, and the trend rules that decide when a slow decline becomes a warning you cannot miss. The daily review is the least glamorous of PupPal's five skills, but it may be the most important — it is the part of the system that notices what no single check-in, and no single owner, can see on any given day.

Why a Single Photo Is Not Enough

Photo check-ins are the heartbeat of PupPal. When you snap a photo, the Cloudflare Worker stores it in R2 and fires the photo webhook; the photo-analyze skill runs one vision call that extracts posture (V2), expression (V4), and environment (V5), then a pure-Python rule engine checks for anomalies (C7). The result is a rich snapshot: the dog was lying down, calm, indoors, no hazards visible, no anomalies flagged. The app shows you that analysis in seconds.

But a snapshot has a fundamental limitation: it has no memory. The dog lying down at 14:32 could be resting after a great walk, or it could be the third afternoon in a row spent listless on the couch. The photo itself cannot distinguish those two realities. Only comparison over time can — and comparison over time is what humans are surprisingly bad at.

The drift problem

Gradual decline is invisible by design. A happiness score that moves from 0.72 to 0.68 to 0.64 across three weeks never triggers the alarm a single dramatic drop would, because each day looks "normal enough." Health problems in dogs rarely announce themselves with one dramatic event; they usually arrive as a slow accumulation of small changes — a little less energy, a little less interest in food, a little more time lying down. By the time an owner notices the pattern with their own eyes, the change has often been underway for weeks.

Memory makes it worse. The psychologist Elizabeth Loftus would not be surprised to hear an owner swear the dog "hasn't been right since Tuesday," when the data shows the decline started ten days earlier. Anchoring, recency bias, and plain busyness all distort our sense of what "normal" is. The pet daily health review exists because a numeric baseline does not lie and does not forget: it stores every day's scores in memory and compares each new day against the previous seven, mechanically and without mood.

What the review adds to each photo

The review turns a stream of snapshots into a longitudinal record:

  • Baseline context. Today's happiness, energy, and health score are compared against 7-day means, producing deltas and a stable/improving/declining verdict.
  • Feeding confirmation. Photos are scanned for eating and drinking postures, and today's detected meals are compared with the pet's usual count.
  • Activity comparison. High-activity moments (running, jumping, playing) and low-activity moments (lying, sleeping) are scored and compared with the 7-day average.
  • Trend detection. The last seven days of history are checked against explicit rules for consecutive decline — the early warning layer that single-photo analysis cannot provide.

The C7 anomaly check already flags acute problems at photo time; the daily review is its chronic counterpart. Where C7 asks "is something wrong right now," the review asks "is something slowly going wrong."

The 21:00 Cron: How the Nightly Ritual Starts

The review does not wait for the owner to remember to run it. It is triggered by a declarative cron job shipped inside the Hermes profile distribution, defined in hermes/cron/daily-review.json. The whole manifest is seven lines, and every field is a design decision.

{
  "schedule": "0 21 * * *",
  "prompt": "运行 daily-review skill:读取今日所有打卡记录,做 S1 基线对比、S3 饭食确认、S4 活动量对比,计算健康评分,检测异常趋势。如果有异常,推送到微信。如果今天无打卡,静默跳过。",
  "skills": ["daily-review"],
  "deliver": "origin",
  "enabled_toolsets": ["memory", "terminal", "file", "messaging", "vision", "web"],
  "name": "Puppal 每日回顾"
}

Why 21:00

The schedule 0 21 * * * runs the review at nine in the evening, local time, every day. The choice of 21:00 is deliberate on both sides of the clock. It is late enough that the day is essentially over — the owner has finished work, dinner is done, and the last check-in of the day has almost certainly been taken. It is early enough that a warning still has the whole evening to be acted on: if the review flags something, the owner can look at the pet, call the vet in the morning, or adjust the care plan before bed.

The timing also sequences the two daily crons cleanly. The review runs at 21:00 and writes its results into memory; the daily-voice cron at 09:00 reads those results the next morning to write the pet's first-person message. The voice is the emotional front end of the review's analytical back end — the dog literally "talks about" yesterday's trends the next morning.

The enabled toolsets

The manifest opens five toolsets: memory, terminal, file, messaging, vision, and web. Memory and terminal are the workhorses — memory reads the day's photos and state, terminal runs the trend-calculation script. File is available for staging output. Messaging is the push channel when a warning fires. Vision and web are available as fallbacks in case the skill needs to re-analyze an image or consult external context; the skill itself is designed to run without them on a normal night.

The skip rules

Not every night produces a report. The skill's "When to Use" section is explicit about the two degenerate cases:

  • No photos today. The review reads the day's check-in records, finds none, and exits quietly with {"status": "skipped", "reason": "no photos today"}. No report, no push, no noise. A missing day is treated as a missing day — not as a health signal.
  • Fewer than two photos. With a single photo there is not enough signal to compare postures or feeding scenes reliably. The skill generates a lightweight report instead of running the full trend analysis, and it labels the analysis confidence as low.

This restraint is what keeps the review trustworthy. A system that produces confident trend analysis from one photo would cry wolf; the review would rather stay silent than guess.

The Three Analyses: S1, S3, S4

Once the cron fires, the skill reads three memory keys — the day's photos (dog:{dog_id}:photos), the pet's current state (dog:{dog_id}:state), and the historical stats (dog:{dog_id}:stats) — and hands them to a pure-Python script, scripts/trend_calc.py. The script needs no external API calls; it runs entirely on the data the photo pipeline already produced. It returns three analysis blocks that map to the three S-numbered checks in the product spec.

S1: Baseline comparison

The baseline comparison answers the question "is today normal?" It extracts the day's expressions from the V4 analysis of every photo, maps each emotion to a numeric score, and weights it by the vision model's confidence. The emotion-to-score table in the script is deliberately simple:

| Emotion | Score | |---|---| | happy | 0.9 | | playful | 0.85 | | curious | 0.7 | | calm | 0.65 | | alert | 0.6 | | neutral | 0.5 | | tired | 0.35 | | anxious | 0.2 | | sad | 0.1 |

Today's happiness is the mean of those confidence-weighted scores. Energy comes from the pet state, which the photo pipeline updates with an exponential moving average that gives the newest observation a 0.3 weight. Then the script pulls the last seven days from history (if at least three days exist) and computes deltas: happiness minus the 7-day mean, energy minus the 7-day mean, health minus the 7-day mean.

Two thresholds do the interesting work. A delta larger than 0.1 in happiness or energy is flagged as a significant change — the report will literally say "energy is down 12% versus the 7-day average." For the overall verdict, a health delta below -0.15 means "declining," above +0.05 means "improving," and anything between is "stable." These are deliberately asymmetric: it takes a smaller improvement to call a trend good than a decline to call it bad, because the cost of a missed decline is much higher than the cost of an under-reported improvement.

S3: Feeding confirmation

The feeding check is almost embarrassingly simple, and that is its strength. The script counts today's photos whose V2 posture is eating or drinking and calls that the number of detected meals. It then compares that count with the pet's usual meal count — which defaults to 2, but is refined to the mean of the last seven days once enough history exists.

The comparison is a ratio test:

  • 0 meals → unable_to_detect (the camera may simply have missed mealtimes)
  • Fewer than half the usual count → less
  • More than 1.5x the usual count → more
  • Anything in between → normal

Note the honesty built into the unable_to_detect case. PupPal does not claim the dog did not eat just because no photo caught an eating posture; it only reports what the photos show. The assessment text says exactly that: "today only detected N eating scenes (usually M), which could mean photos were missed or feeding was light." The review reports uncertainty instead of manufacturing certainty — the same discipline the photo pipeline applies when vision confidence drops below 0.3.

S4: Activity comparison

The activity check quantifies how the dog spent its day. The script classifies each photo by posture: running, jumping, and playing count as high-activity moments; lying and sleeping count as low-activity moments. Every photo contributes a weighted value to the day's activity score:

activity_score = (high_count * 1.0 + neutral_count * 0.5 + low_count * 0.1) / total

A day of pure play scores 1.0; a day spent entirely sleeping scores 0.1. That score is compared against the 7-day average, and a delta beyond ±0.15 flips the trend to increasing or declining. The report then renders it in human terms: "today's activity is 19% below the 7-day average, mostly rest" — the kind of sentence that would take an owner weeks of careful observation to produce on their own.

The three analyses are complementary: S1 tracks mood, S3 tracks food, S4 tracks movement. Together they cover the three axes on which a dog's wellbeing visibly declines — and they feed a single number that summarizes all of them.

The Health Score: One Number, Five Signals

The review's centerpiece is the health score, a 0-to-1 number computed every night from five weighted signals:

health_score =
    happiness * 0.25
  + (1 - anomaly_score_avg) * 0.30
  + energy * 0.15
  + feeding_normal_score * 0.15
  + activity_vs_baseline_score * 0.15

Every weight is a statement about what matters most:

  • 30% — absence of anomalies. The inverse of the day's average C7 anomaly score dominates the formula. Acute problems outrank everything else: a dog with visible skin issues, environmental hazards, or behavioral red flags gets its score pulled down hard, no matter how happy it looks.
  • 25% — happiness. The confidence-weighted expression score from V4. Mood is the most sensitive early indicator of trouble — dogs in pain get quieter before they get sick.
  • 15% — energy. The EMA-updated energy from the pet state, reflecting the recent trend rather than a single moment.
  • 15% — feeding. How normal today's detected meal count looks against the usual pattern.
  • 15% — activity. The day's activity score versus its 7-day baseline.

The weights also make the score hard to game by any single bad reading. One unhappy photo pulls happiness down but leaves the other four components untouched; the score dips, but it does not crash. That is intentional: the health score is a trend instrument, not a mood ring.

Every night the score, along with the day's happiness, energy, photo count, anomaly count, and the raw S1/S3/S4 blocks, is appended to dog:{dog_id}:stats.history as a dated record. That record is the raw material for everything the review does next — and for the trend rules that decide whether a score is just a number or a warning.

Catching Drift Before It Becomes a Problem

The heart of the review is not the score itself but the rules that interpret its movement. The skill defines five trend conditions, each with a severity level and an action:

| Condition | Severity | Action | |---|---|---| | Health score declining 3 consecutive days, cumulative drop over 0.2 | High | Push alert: "health score has been declining, please pay attention" | | Energy below 0.3 for 3 consecutive days | High | Push alert: "activity level very low for several days" | | Happiness below 0.3 on a single day after a normal day | Medium | Flag in the report | | Anomaly score above 0.5 on a single day | Medium | Flag in the report | | Feeding count abnormal (S3) | Medium | Push a reminder |

The rules are the product of a specific design philosophy: trends, not thresholds. A single bad day is flagged but not alarmed — dogs have off days, and one low happiness reading after a normal day earns a note in the report, not a push notification. It takes consecutive days of decline to trigger the high-severity alert, because consecutive decline is the signature of a real problem rather than noise. Three days of falling health score, three days of energy under 0.3 — those are patterns, and patterns are what the review is built to see.

A week in the life of a declining score

To see the rules work, trace a realistic week. Monday: the owner checks in three times, all happy, energy 0.62, health 0.88 — a normal night, report written to memory, no push. Tuesday: two check-ins, one shows a tired posture, energy 0.58. The review compares against the 7-day mean, notes a -0.08 energy delta inside the 0.1 significance threshold, and stays quiet; one soft day is not yet a signal. Wednesday: energy 0.55, and the feeding check detects only one meal instead of the usual two. Now two independent signals have moved. The review flags the S3 anomaly as medium and pushes a summary — the owner learns that meal detection was low and energy is drifting before anything looks visibly wrong. Thursday: the owner, now paying attention, photographs a bit more; energy stays at 0.55, health has now declined three consecutive days. Cumulative drop crosses 0.2, and the high-severity rule fires: an immediate ⚠️ alert, plus the report's advice to watch tomorrow. Friday: the owner takes the dog to the vet, where a low-grade infection is caught early — treatable with antibiotics instead of a week of suffering. That is the review working as designed: three quiet days, one medium nudge, one loud alert, and a problem caught while it was still cheap to fix.

The push decision table

Once the review knows what it found, it decides how loudly to say it. The skill's thresholds section defines four escalation levels:

| Situation | Action | |---|---| | Everything normal | Write the report to memory only; no push | | Medium-level anomaly | Push a WeChat summary | | High-level anomaly | Push immediately, with the ⚠️ marker in the message | | High anomalies 3 days running | Push plus a recommendation to book a vet appointment |

The restraint is the point. If every night pushed a notification, owners would learn to ignore them; the review only speaks when there is something to say, so when it does speak, it is heard. And the final level — three consecutive days of high alerts — escalates beyond the app entirely, suggesting a veterinary visit. PupPal catches the drift, but it never pretends to be the doctor.

What a report looks like

When WEIXIN_ENABLED=true, a flagged day produces a message like this, translated from the skill's example for a dog named DouDou:

📊 DouDou June 18 daily report

😊 Mood: 0.72 (up +0.03) | ⚡ Energy: 0.58 (down -0.08) | ❤️ Health: 0.88 (stable)

📸 5 check-ins today | 🍖 2 meals (normal) | 🏃 Activity below usual

💬 Today's voice: "I was a bit lazy today, spent most of the time
sleeping on the sofa. But I was very happy when we went out this
afternoon!"

⚠️ Needs attention: energy has been declining for 2 days. If it
keeps declining tomorrow, please take note.

That last line is the review doing its job: it does not wait for a third day to warn, it tells the owner what to watch for tomorrow. The photo pipeline's response already includes a trend_brief field with exactly this shape of sentence ("energy declining 2 days in a row, 0.68 to 0.58"), so the review's trend language is consistent with what owners already see at check-in time. For a deeper look at the acute-detection layer that feeds the anomaly score into this formula, see our breakdown of C7 anomaly detection.

What Happens Behind the Scenes: Memory and Honcho

The review is data plumbing as much as it is analysis, and the plumbing has to be honest about its own limits. The skill's Pitfalls section documents four constraints that shape what the review can and cannot claim.

New pets have no baseline

For the first seven days of a pet's life in PupPal, there is not enough history to compute a baseline. The script requires at least three days of history before it will compare against a 7-day mean; before that, the review performs only today's analysis and explicitly skips trend alerts. This is the same rule the C7 anomaly check follows for behavioral anomalies — no baseline, no comparison, no false alarms.

Thin days get low confidence

With a single photo, posture and feeding analysis are statistically meaningless, so the report labels its confidence as low and skips the full trend analysis. The review prefers an honest "we could not tell" over a confident guess, because a confident guess is how warning systems lose their credibility.

Memory reads are bounded

Honcho memory can be slow when asked for long histories, so the skill limits reads to the last 30 days, with older data sampled rather than fully loaded. The trend calculations themselves only use the last seven days anyway, so the 30-day window is a comfortable buffer for the moving averages.

Push is optional

The cron environment may not have WeChat delivery configured. The skill checks WEIXIN_ENABLED before attempting any push and writes the report to memory regardless — the analysis is never lost just because the delivery channel is unavailable.

These constraints are why the review earns trust. It never overstates its data, it never fabricates a baseline, and it always stores its work even when it cannot push it.

Where the Review Goes Next

The nightly review is not an endpoint; it is the analytical backbone for the rest of the product. The 09:00 daily voice reads the review's output to write the pet's first-person morning message, so a quiet decline shows up in the pet's own worried tone the next morning — an emotional nudge that a dashboard could never match. During a foster care session, the care-monitor skill tightens anomaly thresholds by 40%, so the trend layer becomes even more sensitive when someone other than the owner is watching. And on the roadmap — pet check-in, foster proxy check-in, adoption — the review's history is the natural foundation: a rescued pet's first two weeks of trend data is precisely the evidence an adoption match should consider.

The pet daily health review is the quiet part of PupPal, and that is exactly why it matters. Photo check-ins are the heartbeat; the review is the doctor's rounds. Every night at 21:00 it reads the day's photos, compares them with the seven days before, computes a health score from five weighted signals, and decides whether anything is drifting toward trouble — then it either writes a report in silence or speaks up exactly as loudly as the situation deserves. No dashboard to check, no habit to build, no memory to trust. Just a mechanical, nightly, honest answer to the question every owner asks eventually: is my dog okay, really?

If you have not started a daily photo habit yet, tonight is a good night to begin. The review needs history to do its best work — the first seven days build the baseline, and every day after that makes the trends sharper. Snap the photo, share the care, and let the 21:00 review keep watch while you sleep.