C7 Anomaly Detection: How PupPal Flags Pet Health Problems from a Single Photo

2026-08-08

C7 Anomaly Detection: How PupPal Flags Pet Health Problems from a Single Photo

Every PupPal photo check-in runs two pipelines in parallel. While the vision model reads your pet's posture, expression, and surroundings, a second, deliberately boring piece of machinery scans the same result for anything that looks wrong. It is called C7 anomaly detection, and it is the quiet safety net of the whole product: no extra API calls, no extra latency, no extra cost — just a deterministic Python rule engine that decides whether a photo contains a health signal worth escalating.

In the previous post in this series, we walked through the complete photo check-in pipeline: image preparation, the one-shot agent session, the versioned V2/V4/V5/C7 analysis contract, and the WebSocket delivery into the local database. This post zooms in on the C7 half of that contract. We will look at why PupPal chose a deterministic rule engine instead of another model, exactly what the four detectors check, how the anomaly score is computed, how escalations are triggered, and how the whole thing behaves when the inputs are bad, blurry, or just too new.

Why C7 anomaly detection is a rule engine, not another model

The most important design decision in C7 is also the least glamorous: it is not a machine learning model at all. C7 lives in a single file, anomaly_check.py, in the photo-analyze skill, and it is pure Python with zero external dependencies. It takes the vision analysis result and the pet's recent photo history, applies keyword rules and simple statistics, and returns a verdict.

That choice deserves explanation, because it is counterintuitive. If you already have a vision model reading the photo, why not ask it to also judge whether something is wrong? The answer has four parts:

Determinism. A vision model may describe the same photo slightly differently on two runs — temperature, sampling, and prompt phrasing all introduce variation. C7 maps the description to a fixed outcome: the same photo always produces the same anomaly verdict. When a health flag reaches an owner, you want to be able to say "here is exactly why," and to have that reasoning be reproducible. A keyword rule is reproducible by construction.

Cost. Vision calls are the expensive part of a check-in. Adding a second model call to double-check the first one doubles the marginal cost of every photo, forever. C7 runs locally in milliseconds, on the same result that was already paid for. The marginal cost of anomaly detection is zero.

Testability. The engine's behavior is specified as a table of keywords and thresholds, which means it can be unit-tested like any other code. PupPal can add a keyword, tighten a threshold, or fix a false positive with a one-line change and a test — no retraining, no prompt engineering, no regression risk from a model update.

Auditability. When an alert is escalated, the description string tells you the exact signal that triggered it: which keyword matched, or which statistical deviation was detected. There is no black box between the photo and the alert.

None of this means vision is untrustworthy — it means PupPal deliberately separates perception (what is in the photo, handled by the model) from judgment (whether it matters, handled by rules). Perception is fuzzy by nature; judgment should not be.

Where C7 fits in the check-in pipeline

The photo-analyze skill defines the full flow. When an owner takes a photo, the app uploads it and a Cloudflare Worker webhook triggers the skill:

App photo → Cloudflare Worker → Webhook POST /puppal-photo
  → photo-analyze skill
    → Vision API: V2 (posture) + V4 (expression) + V5 (environment) — one call
    → Python script: C7 (anomaly detection) — local computation
    → Memory: update dog_state + append photo_history
  → analysis result → Worker → App

C7 sits between the vision call and the memory update. It consumes the structured vision output — it never sees the raw image — plus the pet's recent history, and it produces the c7_anomalies list, the anomaly_score, and the needs_owner_attention flag that the rest of the system acts on.

There is an important variant: the App Check-in mode. When a check-in happens from inside the app on the local engine track, the app performs a single vision call that merges V2 + V4 + V5, and C7 is skipped — the comment in the skill is explicit that C7's local computation is reserved for the full webhook pipeline. Habit classification in that mode is done client-side with keyword matching on the returned tags, at zero extra token cost. This is the token-optimization strategy that runs through the entire product: every redundant model call is a candidate for elimination, and the parts that can be pure computation are kept pure.

The four C7 anomaly detection detectors

C7 runs four detectors in sequence, each producing zero or more anomalies tagged with a type and a severity. Let's look at each one, with the actual keyword tables from the source.

Posture detection

The posture detector reads the v2_posture block — the position and its free-text notes — and matches against two keyword lists of very different weights.

The first list, POSTURE_EMERGENCY_KEYWORDS, is short and alarming: collapsed, collapsing, unconscious, seizure, convulsing, unable to stand, cannot stand, limping severely, dragging, paralyzed, not moving, unresponsive. Any one of these in the notes produces a critical anomaly with the recommendation to check the photo immediately and contact a vet if necessary. One critical flag is enough — the loop breaks after the first match, because a single emergency signal already changes the priority of everything that follows.

The second list, POSTURE_CONCERN_KEYWORDS, is for softer signals: limping, hunched, trembling, shaking, stiff, unusual posture, abnormal stance, head tilted, circling, pacing restlessly, excessive panting position. These produce a medium anomaly — noted in the daily review, not alerted — and only if no critical flag was already raised.

There is also a third, quieter rule: if the model could not classify the position at all (the other bucket) but was highly confident in that classification (confidence above 0.5), C7 emits a low anomaly — a recognition failure that is worth recording, with a gentle suggestion that the photo angle might be the problem.

The design here is worth noting: severity is not a free choice of the engine author, it is bound to the actionability of the signal. Emergency keywords describe states where minutes matter. Concern keywords describe states where observation matters. Unrecognizable poses describe states where the system's own uncertainty matters. Each level maps to a different human response.

Appearance detection

The appearance detector searches for visible health signals in the notes of both the posture and the expression blocks. Its keyword list is long and specific: blood, bleeding, wound, injury, cut, swollen, swelling, rash, redness, inflammation, lump, discharge, vomit, diarrhea, feces, urine abnormal, hair loss, bald spot, scratching excessively, eye discharge, nose discharge, drooling excessively.

Any hit is a high anomaly — the highest non-critical severity — because appearance signals are exactly what an owner might miss in a cursory glance but a model reading the notes can catch reliably. The recommendation is measured, not alarmist: check the photo to confirm, and seek medical advice if it persists.

Environment hazard detection

The environment detector is the most literal of the four. It reads the v5_environment block and its hazards_visible list — hazards the vision model explicitly saw in the frame — and emits a high anomaly for each one, with the recommendation to remove the hazard immediately.

But it also does something subtler. It matches the environment notes against HAZARD_ALERT_KEYWORDS, a toxicity list: chocolate, grapes, raisins, onion, garlic, xylitol, medication, pills, cleaning product, chemical, poison, sharp object, broken glass, electrical cord chewed, toxic plant, small object swallowed.

Each of these is a critical anomaly — not because the system knows the pet ingested something, but because the possibility of ingestion is time-sensitive. Chocolate and xylitol poisoning in dogs escalate within hours; waiting for the next check-in to confirm is not an option. This is the detector that turns a photo of a dog sitting next to an open wrapper into an immediate owner alert.

Pattern deviation detection

The fourth detector is the only one that needs history, and it is the one that makes C7 feel like it knows your pet. check_pattern_deviation compares the current posture against the pet's own recent behavior.

Two rules gate it carefully. First, if the pet has fewer than seven historical records, the detector refuses to run at all — there is no baseline to compare against, and the code would rather stay silent than guess. New pets get a grace period by construction. Second, when it does run, it looks at the last 14 check-ins and counts how often the current position appears. If a position shows up in fewer than 10% of those records — say a dog that has been photographed lying down for two weeks suddenly appears standing, repeatedly — C7 emits a medium anomaly.

The recommendation text is refreshingly sensible about false positives: "observe whether it persists; a one-off may just be normal play." The detector is explicitly conservative because pattern deviation is the weakest signal of the four — it can be caused by a new camera angle, a new room, or a new toy, all of which are perfectly healthy. C7's philosophy, repeated across every detector, is that an under-claim is recoverable and a false alarm erodes trust. Alerts that fire constantly stop being read, so the engine would rather flag a deviation for the daily review than push it to the owner's phone.

Scoring: from anomalies to a single number

Every anomaly carries a severity — low, medium, high, or critical — and the severity levels have a dual life: they define both the action and the math.

On the action side, the levels map to escalating responses:

| Severity | Action | |---|---| | low | Record only | | medium | Flag in the daily review | | high | Notify the owner | | critical | Alert immediately |

On the math side, the levels map to weights: 0.1 for low, 0.3 for medium, 0.6 for high, 0.9 for critical. The anomaly_score is the sum of the weights of all anomalies in the check-in, capped at 1.0 — which means the score is a density measure, not an average. One critical anomaly (0.9) dominates a stack of low ones, which is exactly the semantics you want: a single emergency signal should not be diluted by a pile of minor notes.

The score is rounded to two decimals and shipped with the result, and on the Flutter side the same number is clamped between 0 and 1 again when it is parsed, so a misbehaving backend cannot produce an out-of-range value that breaks the UI.

Alongside the score, needs_owner_attention is computed by a single rule: true if any anomaly has severity high or critical, false otherwise. It is deliberately not a threshold on the score — a check-in with one high anomaly and a score of 0.6 should trigger the same owner notification pathway as one with three medium anomalies and a score of 0.9. The boolean is the UI's contract; the score is the analyst's.

Escalation: what happens when the score crosses a line

The skill's threshold table turns the score into concrete actions:

  • anomaly_score > 0.5 — the returned message is marked with a ⚠️ and the owner is pointed back at the photo to look for themselves.
  • anomaly_score > 0.7 — an additional standalone WeChat alert is pushed, when the channel is enabled. This is the "this is more than a note" line.
  • Any critical anomaly — pushed immediately, never waiting for the result message to be assembled. The photo may still be uploading; the alert does not care.
  • A single-day health_score drop greater than 0.15 — marked 📉 in the message as a trend worth watching. This is the one threshold that lives outside the photo itself, comparing the pet's rolling health state between days.

There are two interesting properties here. First, the critical path does not wait for anything — it is a separate branch that fires the moment the anomaly list is computed. Second, the thresholds are lines, not curves: 0.49 and 0.5 are different worlds, and that is fine, because the score is deterministic. The same photo always lands on the same side of the line, so the behavior is predictable to both the system and, over time, the owner.

Foster care: when the bar gets stricter on purpose

C7's most deliberate deviation from its own conservatism is the foster care mode. When a check-in comes from a caregiver — the payload's source field is caregiver rather than owner — the same engine runs with every threshold lowered by 40%.

The reasoning is a direct product of the sharing model described in the first post in this series: care codes and PINs let a friend or family member care for your pet with no account at all. But a caregiver does not have the owner's years of context. They do not know whether "sleeping in the afternoon" is normal for this particular dog, or whether that small limp has been there for a week. When someone else is watching your pet, you want to hear about anything unusual sooner rather than later, so the ⚠️ line, the WeChat line, and the pattern-deviation frequency threshold all shift down by 40%.

The trade-off is accepted explicitly: more noise, less risk. A foster check-in that would merely annotate a ⚠️ becomes an active alert, and that is the correct failure mode for the scenario. The care-monitor skill carries this tightening through the rest of the care session, so the entire monitoring posture is different during foster periods — not just the anomaly thresholds but the cadence and sensitivity of review.

The contract on the client side

C7's output is not a free-form paragraph; it is a versioned part of the analysis contract, and the Flutter client models it strictly. In photo_analysis_result.dart, the PhotoAnalysisResult class carries c7Anomalies as a list of Anomaly objects, each with four fields:

  • type — one of posture, appearance, environment, behavioral, or other;
  • severity — one of low, medium, high, critical, or other;
  • description — the human-readable reason;
  • recommendation — the concrete next step.

The enums are parsed defensively: any unknown value falls back to other, and because the LLM often returns Chinese labels, the parser accepts both the enum name and its Chinese label. The anomalyScore is clamped to 0–1 on parse, and needsOwnerAttention is a strict boolean.

Versioning is the contract's most important property. C7 is named as a version, like V2 (posture) and V4 (expression) and V5 (environment), precisely so the schema can evolve without breaking older clients. When the anomaly vocabulary changes — new types, new severity semantics — the version bumps, and clients that understand the old version can keep working against the new one. A health-safety contract is the last thing you want to change silently.

The UI consumes the boolean and the score directly: needsOwnerAttention drives the alert presentation, and the message the owner sees is built from the analysis summary, with the ⚠️ or 📉 markers appended by the escalation rules.

Failure handling: what C7 does when things go wrong

A health-monitoring system earns its keep in the failure cases, and C7 has a well-defined posture for each one:

Malformed input. If the vision result is not valid JSON, the script does not crash the check-in. It exits cleanly with an error field, empty anomalies, a zero score, and needs_owner_attention: false. The check-in survives; only the anomaly layer is skipped.

Low confidence. If the vision model returns a confidence below 0.3 for any analysis block, the skill marks that block uncertain rather than inventing an analysis. The skill's pitfall list is explicit: "never fabricate uncertain analysis." Blurry photos, dark rooms, and partial occlusion produce honest uncertainty, not confident nonsense.

Timeouts. Vision calls can stall on large photos. The skill sets a 30-second timeout; on timeout, the vision analysis is skipped and only C7 runs — on whatever context is available. Degradation is partial and orderly, not total.

No baseline. New pets and pets with fewer than seven historical records do not trigger pattern-deviation detection at all. The engine refuses to guess about "normal" when it has no data, and the skill's pitfall notes the same for the first week of behavior monitoring.

Memory write failures. If the memory backend is unreachable when the skill tries to update the pet's state, the skill checks the connection before writing — the state update is not allowed to silently half-succeed.

Every one of these cases shares the same philosophy: when in doubt, do less, say so honestly, and let a human look. C7 would rather under-claim on a blurry photo than cry wolf, because the cost of a false alarm is not just a notification — it is the erosion of trust that makes the next, real alarm get ignored.

The state that accumulates: memory as the pet's health record

C7's output does not disappear after the check-in. The skill writes the analysis back into the pet's persistent state in memory, and that state is what makes future check-ins smarter.

The pet's dog_state tracks happiness, energy, health_score, and counters like photo_count_today. Updates are smoothed with an exponential moving average that gives the newest observation 30% weight: new_happiness = current_happiness * 0.7 + v4_score * 0.3. The EMA structure means a single bad day moves the needle slightly, not catastrophically — a pet does not become "unhappy" because of one check-in, and a genuinely sustained change accumulates visibly.

Every photo record — with its V2, V4, V5 blocks, its C7 anomalies, and its caregiver notes — is appended to the pet's photo history. That history is the input to the pattern-deviation detector, the baseline for the daily review's trend calculations, and ultimately the raw material for catching health drift before it becomes a problem. C7 does not just flag today's photo; it enriches the record that tomorrow's photo will be judged against.

Why this design holds together

C7 anomaly detection is a case study in restraint. It uses no model where a rule works, no history where a keyword suffices, no alert where a note suffices, and no guess where the data is insufficient. Every design decision — deterministic rules over probabilistic judgment, severity bound to actionability, a capped weighted score, conservative pattern deviation, a 40% tighter foster mode, graceful degradation on every failure path — pushes in the same direction: a health signal that owners can trust because it is reproducible, explainable, and honest about its own uncertainty.

The result is that a single photo, taken in seconds, produces a health screening with the rigor of a checklist: posture emergencies, appearance signals, environmental toxins, and behavioral deviation are each checked against explicit rules, scored, escalated, and stored in the pet's permanent record — at zero marginal cost on top of the vision call that was already happening.

If you want the full picture from the beginning, start with what PupPal is, then read the photo check-in deep dive for the end-to-end pipeline this detector plugs into. And if the C7 approach — deterministic judgment layered on probabilistic perception — is the part you find interesting, the next post in this series looks at the care-monitor skill and how foster care changes the entire monitoring posture, not just the thresholds.