Foster Care Monitoring: Tighter Thresholds, Fewer Assumptions — PupPal

2026-08-10

Foster Care Monitoring: Tighter Thresholds, Fewer Assumptions — PupPal

Every pet owner who has ever left a pet with someone else knows the feeling: the pet is safe, the caregiver is trusted, and yet the questions keep coming. Is the dog eating? Is the cat hiding? Did anyone notice the limp that seemed to be getting better? The moment a pet enters foster care — whether that means a weekend with a neighbor, a week with family, or a longer placement with a foster home — the owner loses their most reliable monitoring instrument: their own eyes. PupPal's answer is a mode of foster care monitoring that deliberately watches harder when someone else is watching. Every photo a caregiver uploads is run through the same vision pipeline, but the rules that decide whether a result deserves an alert are stricter, the checks are more numerous, and the owner is notified instantly instead of at the next scheduled review.

This post explains why foster care monitoring cannot simply reuse the thresholds an owner's own check-ins use, how the care-monitor skill implements the difference inside Hermes Agent, what the 40% threshold reduction actually changes in practice, and why the design deliberately accepts a higher false-alarm rate. It builds directly on the earlier deep dives into the photo check-in pipeline and the C7 anomaly engine, and it assumes the sharing model described in the care codes post. If you have not read those, the short version is: a photo is a check-in, sharing a care code is how you let someone else check in on your behalf, and every check-in photo flows through a Cloudflare Worker into Hermes Agent for analysis. This post is about what happens to that analysis when the person holding the phone is not the owner.

The Problem: Someone Else Is Watching Your Pet

Start with the asymmetry that defines foster care. When the owner takes a photo, they bring a lifetime of context to the result. They know that the dog lies like that after a long walk, that the cat squints when the sun hits the window at 4pm, that the low-energy afternoon is normal for a Tuesday. The anomaly engine does not need to compensate for missing context, because the owner is the context. If the analysis says nothing unusual, the owner can confirm it with their own judgment in seconds.

A caregiver has none of that. They may be meeting the pet for the first time, or the second. They do not know whether the dog always trembles slightly when settling down, or whether the cat's hiding behavior under the sofa is her personality or a signal of distress. The caregiver cannot distinguish a normal quirk from an early symptom, which means the system must do that work for them. That is the core argument for a separate monitoring mode: the same photo, analyzed by the same model, deserves a different judgment depending on who took it and when.

There is also a second asymmetry, and it is about attention rather than knowledge. During a normal day, a photo check-in is one input among many; the owner sees the pet continuously and the photo is a periodic confirmation. During foster care, the caregiver's photos are often the owner's only window into the pet's condition. A weekend away with one photo per meal means the owner sees maybe six data points across 48 hours. Every single one of them must count. Missing a signal in one of those six photos is not like missing a signal in a daily routine of ten; there is no next photo to correct the error, and there may be no one with the knowledge to notice.

The design conclusion is uncomfortable but unavoidable: foster care monitoring must be optimized for catching problems early, not for avoiding false alarms. In statistical terms, the mode trades precision for recall — deliberately, and with the owner's informed consent baked into the product's behavior.

What Changes When a Foster Session Begins

Before any special monitoring can happen, the system has to know that a foster period exists at all. That is the job of the care session, a first-class object that the Cloudflare Worker and Hermes Agent share.

The care session lifecycle

The owner starts a session through the care code flow described in the care codes post: the Worker generates a code in the PUPPY-XXXX format plus a PIN, and the owner shares them out of band. Underneath that user-facing flow, the Worker fires the puppal-care-create webhook into Hermes Agent. The payload is small and specific:

{
  "dog_id": "doudou",
  "care_session_id": "care_abc123",
  "start_time": "2026-06-18T10:00:00Z",
  "end_time": "2026-06-20T18:00:00Z"
}

The session carries an optional end_time. If it is set, the session expires automatically; if it is not, the session lives until the owner ends it. Either way, the ending path is the same: the Worker fires puppal-care-end with a reason of manual or expired, the agent updates the session record's status to ended, and the care code is invalidated immediately. Session state lives in Memory at dog:{dog_id}:care_sessions.{care_session_id} — a per-session record that the monitoring skill reads on every caregiver photo.

The care handbook: context for a stranger

The puppal-care-create webhook is handled by the care-handbook skill, which generates a care handbook: an LLM-written, pet-specific set of instructions covering feeding, behavior, quirks, and anything the owner wants the caregiver to know. The handbook is the caregiver's substitute for the owner's context. It is shown when the caregiver enters the code and PIN, and it is the first thing a caregiver reads before taking their first photo.

The handbook matters for monitoring for a subtle reason: it is the only place where the caregiver's expectations are set. A caregiver who has been told "she hides for the first day, this is normal" will not panic when the first photo shows the cat under the sofa; a caregiver who has not been told will interpret the same photo through anxiety. By pairing an AI-generated handbook with AI-driven monitoring, the system aligns the human and the algorithm on the same expectation set — which is precisely why the monitoring mode can be stricter without becoming useless.

One photo pipeline, two judgment modes

Every photo, owner or caregiver, arrives through the same POST /webhooks/puppal-photo webhook. The routing decision happens inside the agent: when the request body carries source="caregiver" along with a care_session_id, the photo is handed to the care-monitor skill instead of the standard photo-analyze skill. Nothing else changes upstream — same R2 storage, same Worker relay, same Vision call. The divergence is entirely downstream, in the rules that interpret the analysis result.

Why Foster Care Monitoring Needs Its Own Rules

The care-monitor skill makes its purpose explicit in its own documentation: caregivers are not familiar with the pet, and they are easy to miss subtle abnormal signals. Care Monitor is therefore deliberately more "sensitive" and more "verbose" than owner-mode monitoring — the skill's own language says it would rather report once too often than miss once. That stance is the design philosophy, and it shows up in three concrete engineering decisions.

The first decision: route, don't fork

Care monitoring does not reimplement photo analysis. Its procedure starts by validating the session (care_session.status == "active", otherwise the photo is rejected), then runs the same V2/V4/V5/C7 analysis as photo-analyze — the Vision model classifies posture, expression, and environment in one call, and the Python rule engine computes anomalies — with two differences: the result is marked with source="caregiver", and the care_session_id is attached to every recorded outcome. Reusing the pipeline guarantees that the two modes never drift apart in what they perceive; they only differ in how they judge.

The second decision: tighten what the owner mode already checks

The standard mode has two push thresholds that govern how aggressively results become alerts. Care mode lowers both by 40%, and also lowers the bar on the emotional-drift check that watches for consecutive low-mood photos. The exact numbers are in the next section. The rationale is simple: in owner mode, a borderline signal can wait for the next check-in or the nightly review; in care mode, a borderline signal may be all the owner gets for hours.

The third decision: add checks that only make sense during foster care

A few detections have no meaning in owner mode. "The environment changed from home to something else" is unremarkable when the owner took the pet to the park themselves, but during foster care it is a signal worth flagging, because the pet may be in an unfamiliar place on top of an unfamiliar routine. Likewise, a caregiver's free-text note ("she didn't eat", "he threw up") is a data source the owner mode never sees, and it deserves immediate, direct escalation. These care-only checks are what make care mode more than just a threshold tweak; they are the part of the design that acknowledges the mode's real job: detecting change from the owner's baseline, not absolute abnormality.

The lesson that generalizes beyond PupPal: monitoring rules should be a function of who is watching and how much context they have, not a fixed property of the signal being monitored. The same pixel data warrants different alarm thresholds in different hands.

The 40% Rule: Thresholds in Care Mode

The centerpiece of the care-monitor skill is a single multiplier: 0.6. Every standard threshold is multiplied by 0.6, which is a 40% reduction, and the constants are written out explicitly rather than computed at runtime, so the two modes can never silently drift. The standard mode defines ANOMALY_PUSH_THRESHOLD = 0.7 (an anomaly score above 0.7 triggers a push) and HEALTH_DROP_THRESHOLD = 0.15 (a health-score drop above 0.15 triggers an alert). Care mode defines the lowered constants CARE_ANOMALY_PUSH_THRESHOLD = 0.42 and CARE_HEALTH_DROP_THRESHOLD = 0.09.

| Condition | Standard mode | Care mode | Difference | |---|---|---|---| | anomaly_score push threshold | 0.7 | 0.42 | ↓ 40% | | health_score drop alert | 0.15 | 0.09 | ↓ 40% | | Consecutive low-mood photos before alert | 5 | 3 | ↓ 40% | | Persistent environment unfamiliarity | not triggered | 3 photos | care only | | Caregiver note contains alert keywords | not triggered | immediate alert | care only |

What the numbers mean in practice

The anomaly score is computed by the C7 engine's calculate_anomaly_score: each detected anomaly contributes a weight by severity, and the score is the sum, clamped to 0–1 in the Flutter client's PhotoAnalysisResult model. In owner mode, a score of 0.5 — say, a medium-severity posture keyword like "trembling" plus a low-severity note — would not push an alert; it would sit in the photo's analysis and be picked up by the nightly daily-review at 21:00. In care mode, that same 0.5 exceeds 0.42, so the owner gets an immediate push with the photo attached.

The health-drop threshold works similarly, but on the longitudinal axis: the engine compares the current photo against the pet's recent state, and a drop of more than 0.09 in care mode (versus 0.15 in standard mode) triggers an alert. The practical effect is that a pet whose health score is drifting downward will surface to the owner roughly twice as early during a foster period — while there is still time to call the caregiver and ask questions, rather than after the nightly review aggregates the trend.

Why 40% and not, say, 50%

The choice of 40% is a judgment call, but it is a principled one. The skill's own documentation is explicit that care mode has a higher false-positive rate by design, and that is acceptable because safety wins over convenience. A 40% reduction keeps the threshold in a band where genuine signals still dominate the alert stream; a smaller reduction would leave too many borderline cases waiting for the nightly review, and a larger reduction would flood the owner with noise and train them to ignore alerts — which is the one failure mode that makes the whole exercise pointless. The 40% figure also creates nice round constants (0.7 × 0.6 = 0.42), which makes the two modes easy to audit: anyone reading the skill can see at a glance that care mode is exactly 60% of standard sensitivity.

Checks That Only Run During Foster Care

Beyond the threshold reduction, care mode runs the standard C7 checks — posture keywords, appearance keywords, environment hazards, and pattern deviation against history — plus five care-specific detections that take the session into account. The anomaly engine is invoked with --mode "care" and --care_start_time, so these checks have the session window available as a boundary.

The first anomaly of the session

The first anomaly of any kind detected after the care session starts is flagged immediately, regardless of its severity score. The reasoning is asymmetric-information again: the very first signal from a foster period is the least expected and the least explainable, so it gets attention by default. A low-severity quirk that the owner would wave off at home becomes a data point the owner should know about during foster care — because the owner is the only person who can decide whether it is a quirk or a symptom.

If the same type of anomaly appears at least twice during the care period, its severity is upgraded. A single low-energy photo might be nothing; two low-energy photos across two days of a three-day foster stay is a trend, and the trend is the thing that matters when you only see six photos in 48 hours. This check is the care-mode analog of the daily-review skill's S1/S3/S4 trend calculations, compressed into the session window instead of the nightly batch.

Caregiver notes as a first-class signal

The caregiver can add a note to a photo, and care mode treats that note as a real signal: if it contains keywords like "not eating", "vomiting", or "diarrhea" — the skill ships Chinese keyword matching (不吃, 吐, 拉稀) and explicitly supports bilingual matching, since caregivers and owners may not share a language — the anomaly is marked high immediately, with no waiting on the score. This closes the loop between the human observation and the machine analysis: the caregiver saw something, the system believes them, and the owner hears about it now.

Environment unfamiliarity

If the Vision model's v5_environment classification says the location type is not the pet's usual home environment — LocationType has values like indoorHome, outdoorPark, outdoorStreet, vehicle, vet — care mode marks a medium-severity anomaly and suggests an adaptation period. The same detection is meaningless in owner mode (the owner took the dog to the park, obviously), but in foster care it is context the owner needs: the pet may be in a new place, and new places change behavior in ways that can look like symptoms.

Persistent low mood

Standard mode alerts after 5 consecutive photos with a non-positive emotion. Care mode alerts after 3. The v4_expression analysis returns one of the Emotion enum values — happy, calm, alert, curious, tired, anxious, sad, playful, neutral — and the engine tracks the run of non-positive results across the session. Three is a deliberately small number for the same reason the anomaly threshold is low: with sparse photos, three consecutive low-mood readings may represent a full day of the pet's experience.

Alerts, Records, and Guardrails

None of the extra sensitivity means anything if the alert does not reach the right person, and none of it is trustworthy without guardrails on the edges. This section covers the alert path, the record trail, and the deliberate constraints that keep care mode from becoming a noise machine.

One recipient: the owner

Care mode pushes alerts only to the owner, never to the caregiver. The push payload carries a title in the form "⚠️ Care alert — {pet_name}", a body with the anomaly score and a suggestion to contact the caregiver to confirm, the photo URL, and the care_session_id. The asymmetry is intentional: the caregiver is already next to the pet and does not need to be told what the photo shows; the owner is far away and needs the earliest possible signal. Also, alerts are pushed immediately — they are not batched into the nightly 21:00 daily-review. Waiting for the daily review would defeat the entire purpose of the stricter thresholds: the value of care-mode sensitivity is that a Tuesday-morning anomaly does not wait until Tuesday night to reach the owner.

The session record

Every caregiver photo is appended to dog:{dog_id}:care_sessions.{care_session_id}.caregiver_photos with its photo_id, timestamp, anomaly_score, whether the owner was alerted (alerted_owner), and an analysis_summary. This record is what turns a foster period into an auditable story. When the owner comes home, they can review exactly what the caregiver's photos showed, what the engine flagged, and what was escalated — the accountability trail that makes sharing care feel like gaining information rather than losing control.

Guardrails: fixed thresholds, low-confidence photos, and honest notes

Three pitfalls in the skill's own documentation show where the design draws its lines:

  • Thresholds are fixed. The skill explicitly warns against raising thresholds after two consecutive false alarms. In owner mode, tuning sensitivity to your own pet is fine; in care mode, consistency is the product. A caregiver's photos are judged by the same rules every time, and the owner knows exactly what those rules are.
  • Bad photos are analyzed, not skipped. Caregivers take worse photos than owners — motion blur, bad light, a pet that will not hold still. When Vision confidence drops below 0.3, the analysis is marked uncertain but still runs. Skipping low-confidence photos would create a blind spot exactly where care mode needs the most coverage.
  • Notes are trusted and bilingual. The keyword matcher supports Chinese and English, because caregivers and owners frequently do not share a language, and a "didn't eat" observation must never be lost in translation.

The guardrails share one theme: care mode optimizes for never missing, and it accepts the cost — more false positives, more uncertain results, more pushes — because the alternative is a system that quietly fails to notice while a stranger is responsible for your pet.

From Foster Care to Adoption: The Road Ahead

Foster care monitoring is not a static feature; it is the middle stage of a roadmap that runs from check-in to adoption. The first stage, pet check-in, is done. The second stage, foster care — with its care codes, care handbook, and the care-monitor skill described here — is shipping. The third stage, adoption, will lean on the exact same machinery, and the current development work in the Puppal repository shows where the foster-care story is heading.

Foster proxy check-in is in active development

The most recent commits in the Puppal repo are about proxy check-in (代照看): letting a caregiver's photos reach the owner's app directly, over WebRTC, rather than only through the cloud relay. The work includes a pure-Dart P2P chunking protocol with a receiver assembler (the inverse of the photo sharder), session orchestration with RTC sending, a care inbox for incoming photos and selection for proxy check-in, presence-based signaling that replaced a 2-second polling loop, a transfer_mode degradation path when P2P is not available, and signaling-room logic that kicks stale connections so a reconnecting member cannot linger as a zombie. The throughline is exactly what foster care monitoring is about: making sure that when the owner is not there, the photos — and the analysis they feed — keep flowing, by whatever path works.

Why foster care monitoring is adoption screening in disguise

An adoption placement is the most extreme foster case: the caregiver may become the permanent owner, the monitoring window may be weeks instead of days, and the stakes of missing a health signal are the highest they ever get. The care session, the caregiver photo record, and the strict thresholds already produce, as a byproduct, the single most useful artifact an adoption decision can have: a longitudinal record of how the pet actually behaved and looked in a home environment, with every anomaly timestamped and every alert documented. A shelter or rescue that can show a foster period's full photo-and-analysis trail is not just monitoring the pet; it is building the case file for adoption.

That is the deeper point of this whole design. Foster care monitoring in PupPal is not a stricter version of the same feature — it is a different product feature with a different job. The job is to hold the owner's place in the loop when the owner cannot be there: to be more sensitive because the watcher has less context, to escalate faster because the photos are rarer, and to record everything because the record is what turns worry into knowledge. If you have a pet, and someone you trust is about to care for it, the care code is how you hand over the responsibility — and foster care monitoring is how you never really let go.

The whole story hangs together: the photo check-in pipeline turns pixels into status, the C7 anomaly engine turns status into signals, and the care code model turns a stranger into a caregiver. Foster care monitoring is what makes the last step safe: the same photo, the same pipeline, and a different, deliberately stricter judgment — 40% tighter thresholds, care-only checks, immediate alerts, and a complete record — so that when your pet is in someone else's hands, your eyes are still on every photo.