Inside a Photo Check-in: What Really Happens When You Snap a Picture of Your Pet
2026-08-07
Inside a Photo Check-in: What Really Happens When You Snap a Picture of Your Pet
In the first post in this series we introduced the core idea behind PupPal: photo = check-in, share = care. A photo of your pet is not just a photo — it is a structured care event that gets analyzed, persisted, and shared with the people who matter. On the surface it looks like magic: point your camera at your dog, and a few seconds later PupPal tells you how they are feeling, whether anything looks off, and what to do about it.
Behind the surface is a surprisingly deep pipeline. This post walks through every stage of a single photo check-in, using the real implementation in the PupPal codebase as our map. By the end you should understand exactly where your photo goes, what the AI actually looks at, how the result is stored, and why the system is built the way it is.
The big picture: one photo, four phases
A single check-in photo travels through four distinct phases:
- Capture and dispatch — the app hands the photo to an AI engine implementation, which decides how and where analysis runs.
- Analysis — an agent session reads the image, applies the pet's profile and accumulated memory, and produces a structured result.
- Delivery and persistence — the result is pushed back to the app over a WebSocket gateway and written to the local database.
- Feedback — the result is rendered, tags are extracted, and any newly learned skills are surfaced to the owner.
Let's look at each phase in detail.
Phase 1: Capture and dispatch
The app defines a single abstraction for all AI work — the AiEngine
interface. It exposes six generation methods (analyzePhoto, dailyReview,
careHandbook, dailyVoice, chat, and classifyActivity), and the rest of
the app never talks to a concrete implementation. That matters because PupPal
currently supports three engine tracks, and switching tracks is a matter
of swapping the implementation, not rewriting the feature layer:
- Local (Rust) — the agent loop runs on-device in a native Rust core, bridged into Flutter via FRB. This is the default path today.
- Hermes (remote) — photos live on the server, so the app sends a
mediaFileIdinstead of image bytes, and the analysis happens in a Hermes Agent session on the backend. - Cloud (managed) — a placeholder for a future managed Worker-AI track.
A photo is passed to the engine as a PhotoRef: either a local file path
(PhotoRef.local) or a remote media file id (PhotoRef.remote). The engine
never has to guess where the image lives — the caller already knows.
The UI layer drives this through a ViewModel that subscribes to an
event stream, not a one-shot future. The engine returns an
AiGeneration handle containing a Stream<AiEvent> and a real cancel()
function. Events flow as AiEventAck (the request was accepted), then a
series of text deltas and tool-progress notifications, and finally
AiEventDone with the complete result. This streaming design is what lets
the UI show live progress — "📸 analyzing photo" — instead of a frozen
spinner. And because cancellation is wired end-to-end (each request carries
its own CancelToken, and even unsubscribing the stream cancels the
request), leaving the screen mid-analysis does not leak an orphaned HTTP
call or a zombie agent session.
Phase 2: Analysis — what the agent actually does
This is where the interesting engineering lives. On the local track, the photo goes through the Rust agent core in a one-shot session:
load image ──▶ check rate limit ──▶ run one-shot agent session ──▶ parse JSON ──▶ persist attempt
Preparing the image
Before anything touches the model, the image is prepared. Files up to 4 MB are sent as-is, encoded as base64. Anything larger is decoded, resized so its long edge is at most 1568 pixels (the recommendation for Anthropic's vision models), re-encoded as JPEG at quality 90, and then sent. This keeps API costs and latency predictable regardless of whether the user's camera produced a 2 MB or a 24 MB file.
The one-shot agent session
The photo is analyzed in a single agent session with three important properties:
- Profile injection — the pet's stored profile is loaded and injected into the system prompt, so the model knows who it is looking at: name, breed, gender, weight. On the remote Hermes track the same thing happens server-side, with the profile serialized as JSON into the system message.
- Skill and memory injection — the session runs with the pet's learned skills and accumulated memory available, not just the raw image. This is what turns a generic vision model into your pet's analyst.
- Three tool iterations — the agent is allowed up to three tool-call rounds inside the session. The comment in the code says it best: "facts discovered during analysis — allergies, quirks, environmental hazards — can be settled into memory while we're here." A photo of your cat next to an open window doesn't just get analyzed; if the agent notices the hazard, it can persist that fact for every future session.
The engine is strict about output. The agent must return its analysis as a
single JSON document, and the app parses that JSON through the same
PhotoAnalysisResult.fromJson contract on every track. If the model returns
malformed JSON, the attempt is recorded as failed and the user gets a
retryable error — the system never silently swallows a bad result.
The analysis contract: V2, V4, V5, C7
The structured result is where PupPal's design philosophy shows. Instead of one blob of free-text commentary, the contract splits the analysis into four versioned sections, each with a strict schema:
V2 — Posture. Position (standing, sitting, lying down, running, jumping, playing, eating, drinking, sleeping, or other), a confidence score, the percentage of the body visible in frame, and notes.
V4 — Expression. The pet's emotion (happy, calm, alert, curious, tired, anxious, sad, playful, neutral), again with confidence, plus a list of concrete indicators the model observed, and notes.
V5 — Environment. Where the pet is: indoor at home, indoor elsewhere, outdoor park, outdoor street, outdoor yard, in a vehicle, at the vet, or other. Crucially it also reports visible hazards, the number of other animals in frame, and the number of people. A check-in from a vet's office or a park carries different context than one from the living room, and the system keeps that context explicit instead of burying it in prose.
C7 — Anomalies. The health-flagging section: a list of anomalies, each
typed (posture, appearance, environment, behavior, other), each carrying a
severity (low, medium, high, critical), a human-readable description, and a
concrete recommendation. Alongside the list is a single anomaly_score
clamped between 0 and 1, and a boolean needs_owner_attention that the UI
can act on immediately.
The names are versioned (V2, V4, V5, C7) deliberately: the contract is a living document, and when the prompt or schema changes, the version bumps with it — old clients and new clients can coexist without guessing what a field means.
Phase 3: Delivery and persistence
Once analysis completes, the result has to get back to the app and into the database. The path it takes is one of the more careful pieces of the design.
The gateway push
The app does not poll. The backend pushes the finished result to the client
over a WebSocket gateway as a JSON-RPC-style event whose method is
puppal.photo.result. The event payload carries:
photo_idanddog_id— which check-in and which pet this is about;analysis— the full structured result described above;dog_state_snapshot— the pet's state as the backend understood it at analysis time, so the app can reconcile anything that changed while the analysis was running;message— the human-facing summary; andphoto_url— the cloud URL of the uploaded image.
Writing to Isar
The handler's job is to turn that event into durable local state, and it does so in three writes:
- Play — the check-in record itself. The AI message becomes the record's description, and the extracted tags are attached. This is the history you scroll through.
- MediaFile — the photo. The cloud URL and upload status are recorded and linked back to the play record, so the app knows the photo is safely stored and can display the remote copy.
- AIAnalysis — the complete, raw analysis result, JSON-encoded, with
its confidence (the
anomaly_score) and the set of analysis sections as keywords. Nothing is summarized away at write time; the full contract is preserved for later use.
Each write happens inside its own isar.writeTxn, and the lookup that
anchors everything (finding the Play by its remote photo id) throws
immediately if the check-in record is missing — a photo result with no
matching check-in is treated as a bug, not silently ignored.
Tags: the result becomes searchable history
The handler also derives tags from the structured analysis: the emotion from V4, the position from V2, and the location type from V5 each become a tag, and the first of those becomes the primary tag. This is a small detail with a big payoff: because tags come from a fixed enum, not free text, the history is consistently filterable. "Happy" and "sitting" and "park" mean the same thing every time they appear, which is exactly what you want when you later ask "how often is my dog happy outdoors?"
Phase 4: The feedback loop — skills unlock
The last phase is the one that makes PupPal feel alive: the check-in doesn't just record, it teaches.
Every pet accumulates a skill list, maintained by the agent core and queried through the Rust bridge. Before the photo analysis starts, the app snapshots the current skill set. When the result comes back, it queries the list again and diffs the two:
List<String> diffNewSkills(Set<String> before, List<PiSkillInfo> after) {
return [
for (final s in after)
if (!before.contains(s.name)) s.name,
];
}
Any skill that appears in the new list and not the old one is newly learned, and the app celebrates it: a banner appears on the check-in screen — scoped to that photo, so replays and old check-ins don't re-trigger it — and the pet's skill set is one click away in the skill tree page, with a badge on the home screen showing how many skills the pet has learned.
This is the loop that gives the product its personality: you check in every day, the agent analyzes, and over time your pet visibly "learns" — the skill tree grows, the analysis gets smarter because memory accumulates, and the owner has a concrete reason to keep the habit alive. It is check-in gamification with a real mechanism behind it: the diff is computed from actual agent state, not faked.
What the user sees — and can correct
Finally, the message itself. The analysis produces a natural-language
summary, but PupPal treats it as editable truth. The result model carries
isEdited and editedMessage fields; when the owner rewrites the message,
the display prefers their version (editedMessage ?? message). The AI is
not infallible, and the product is honest about that — a wrong summary can
be corrected in place, and the correction is what subsequent views show.
Why this design holds together
Stepping back, a few principles explain why the pipeline looks the way it does:
- A strict contract beats free text. The versioned V2/V4/V5/C7 schema is what makes tags consistent, history filterable, and future features (trend calculation, anomaly escalation, foster-care monitoring) possible without re-parsing prose.
- Events, not polls. From the SSE stream that drives the UI to the WebSocket gateway that delivers results, everything is push-based and cancellable. No polling loops, no orphaned requests, no stale snapshots — the dog-state snapshot is delivered with the result precisely so the app can reconcile what changed during analysis.
- The agent has memory and hands. Profile injection, skill injection, and three tool iterations mean a check-in is not a blind model call — it is a small agent that can consult history and persist new facts.
- Everything is recorded, even failures. Failed parses are persisted as failed analysis attempts. Bad results are retryable errors, not silent drops. The audit trail is complete by design.
The next time you snap a photo in PupPal, what you are actually doing is kicking off a small distributed system: image preparation, a rate-limited agent session with memory, a structured health assessment, a push over a gateway, three transactional writes, and a skill diff — all in the seconds it takes to put your phone back in your pocket. And every one of those stages is a decision we can revisit, version, and improve without changing what the check-in means to you.
That's the deep dive. Next in the series: the C7 anomaly detection system — how PupPal decides what counts as a problem, how thresholds work, and when it escalates to a human.