Flutter Rust Bridge: Cutting a 30-Minute Build to Seconds — PupPal
2026-08-09
Flutter Rust Bridge: Cutting a 30-Minute Build to Seconds — PupPal
Flutter gives you a beautiful cross-platform UI in hours. Rust gives you a performance-critical core you can trust. Putting them together is supposed to be the best of both worlds, and with flutter_rust_bridge it genuinely is — until the first time you change one line of business logic in the Rust crate and wait thirty minutes for the Android build to finish. That was the exact situation PupPal hit. PupPal embeds a full AI coding-agent runtime — a Rust port of the pi-mono TypeScript agent, plus its own session bridge — inside a Flutter app, and the compile-time tax of that decision nearly derailed the project.
This post is the optimization playbook we actually executed. It is
grounded in a real codebase: the rust_puppal_api crate that sits at
puppal/rust/ in the PupPal monorepo, the giant pi_agent_rust crate
it depends on, and the cargokit-based rust_builder plugin that wires
them into the Flutter build. Every technique here is one we measured:
shared CARGO_TARGET_DIR, debug profile tuning, release-mode daily
driving, dependency pruning, crate splitting, sccache, and the
discipline of never running flutter clean again. If you maintain a
Flutter app with a Rust core, this is the order in which you should
attack your own build times.
Why Flutter and Rust End Up in the Same App
PupPal is a pet-care app, but its differentiator is not the UI. It is the intelligence layer: every photo a user takes of their pet goes through an analysis pipeline, an AI agent that can reason about the photo, consult care handbooks, run tools, and stream back a natural language status update. That agent is heavy machinery. It has a session model, a tool-calling loop, provider adapters, token accounting, and a streaming event protocol. Writing that from scratch in Dart would mean reimplementing an entire agent runtime; embedding an existing one written in Rust was the only realistic path.
The architecture that resulted looks like this:
puppal/rust/— the bridge craterust_puppal_api, a thinflutter_rust_bridge(FRB) layer that exposes exactly the API the Flutter app needs. Itsrust_inputiscrate::api, its output goes tolib/ruston the Dart side.puppal/packages/pi_agent_rust/— the upstream agent crate, a vendored copy withdefault-features = falseso the CLI/TUI feature set never enters the mobile build.puppal/rust_builder/— a cargokit-based Flutter plugin that drives Cargo from the Flutter build for Android, iOS, Windows, macOS and Linux.
The bridge itself is small but speaks a deliberately simple protocol.
In pi_bridge.rs, a PiSessionConfig struct carries the session
settings from Dart — provider, model, API key, base URL, system prompt,
tool-iteration budget — and a PiBridgeEvent enum streams results back
as single-line JSON strings through an FRB StreamSink. That design
choice (JSON strings instead of FRB custom enums) matters: it keeps the
FFI surface trivially serializable and lets the Rust side evolve its
event vocabulary without regenerating Dart bindings. It is a clean,
minimal seam — and the compile-time monster on the other side of it was
anything but minimal.
The Anatomy of a Flutter Rust Bridge Build
To understand the optimization, you have to understand what actually happens when you press build. A Flutter app with a Rust core does not compile one crate; it compiles a dependency graph, and the shape of that graph decides everything.
rust_puppal_api declares flutter_rust_bridge = "=2.12.0" and
depends on pi via a path dependency:
pi = { path = "../packages/pi_agent_rust", package = "pi_agent_rust",
default-features = false }
That one line is the whole problem in miniature. pi_agent_rust is not
a small library. Its source directory contains files measured in
megabytes: extensions.rs at 2.0 MB, extensions_js.rs at 1.1 MB,
agent.rs and doctor.rs each over 400 KB. It depends on heavyweight
crates — swc_* for JavaScript parsing, ast-grep-* for structural
search, wasmtime, tikv-jemallocator, a TUI stack (crossterm,
bubbletea, glamour), arboard, and more. Its own lib.rs is
explicit about the deal: the crate is primarily the implementation
crate for a CLI binary, and external consumers should treat everything
outside the sdk module as unstable.
So every Flutter build dragged in a CLI agent's entire dependency
universe, because the path dependency did not discriminate. The
default-features = false flag removed the feature-gated extras, but
a large fraction of the heavyweight dependencies are not feature-gated
at all. A one-line change in the bridge could cascade into recompiling
swc and friends. The first time we hit a full rebuild, the Android
build took about thirty minutes on a fast machine. That is not a
workflow; that is a daily penalty that makes you stop touching the
Rust side at all.
The fixes below are ordered by effort-to-reward. Do them in this order and you will see compounding wins after the first two steps.
Step 1: Shared CARGO_TARGET_DIR — Stop flutter clean From Destroying Your Cache
The single most damaging habit in a Flutter+Rust project is flutter clean. It is the documented first resort for every Flutter build
mystery, and it silently deletes the Rust target/ directory if your
plugin's Gradle configuration writes Cargo artifacts under build/.
Deleting the Cargo cache turns the next build into a full cold
rebuild of every dependency — which, for pi_agent_rust's dependency
graph, is the thirty-minute disaster.
The fix is to move the Cargo target directory out of the project tree
entirely. Set a shared CARGO_TARGET_DIR environment variable pointing
at a fast SSD:
# Windows
set CARGO_TARGET_DIR=D:\cargo-target-shared
# macOS / Linux
export CARGO_TARGET_DIR=/Volumes/FastSSD/cargo-target-shared
Three consequences follow. First, flutter clean cannot touch the
cache anymore because the cache is not in the project. Second, the
cache is reused across projects that share Rust dependencies — your
bridge crate and any other Rust tooling on the machine converge on one
artifact pool. Third, incremental compilation actually survives
between builds, because Cargo's fingerprint database is not being
deleted out from under it.
Two refinements. If you use VS Code, mirror the setting in
.vscode/settings.json with "rust-analyzer.cargo.targetDir" so the
language server shares the same cache. And after moving the directory,
delete the old target/ inside the project once so Cargo does not get
confused by two locations — then never delete it again. PupPal also
went one step further in the Android build: the Gradle plugin
(plugin.gradle) was changed to emit Cargo artifacts outside build/
and .gitignore was updated to match, which is what made flutter clean fundamentally incapable of nuking the Rust cache. This one
change alone removed the "half-hour rebuild after every clean" failure
mode permanently.
Step 2: Profile Tuning — The 10-Second Edit That Changes Everything
Cargo profiles are where you get the biggest wins for the smallest
effort, and there is a subtlety that trips up almost everyone: profile
settings in a dependency's own Cargo.toml are ignored when that crate
is compiled as a dependency. The effective profiles are the ones in
the root crate of the build — in PupPal's case, rust_puppal_api.
pi_agent_rust can declare whatever it likes in [profile.dev]; when
it is built as a dependency of the bridge, those declarations do not
apply. You must configure profiles in the crate that Cargo treats as
the workspace root.
For the debug profile, the goal is compile speed and link speed, not runtime speed:
[profile.dev]
opt-level = 0
debug = 1
codegen-units = 16
incremental = true
debug = 1 is the interesting line. Full debug info (debug = 2) is
the default, and it significantly slows linking of large crates because
the linker has to process enormous symbol tables. debug = 1 keeps
only the line-number tables — which is all a debugger actually needs
for stepping and breakpoints — while cutting link time dramatically.
codegen-units = 16 maximizes parallel codegen, and incremental = true makes sure the incremental compilation cache is active.
For the release profile, the priorities flip: you are shipping this code to phones, so size and runtime performance dominate:
[profile.release]
opt-level = "z"
lto = true
codegen-units = 1
strip = true
panic = "abort"
debug = false
debug-assertions = false
overflow-checks = false
PupPal's release profile mirrors what pi_agent_rust itself ships:
opt-level = "z" optimizes for binary size, lto = true with
codegen-units = 1 enables whole-program link-time optimization, and
strip = true removes symbol tables. The measured effect on the
Android shared library was a drop from about 22 MB to the 6–9 MB
range. On top of that, the app's Gradle config was changed to build
release only for arm64-v8a, which further trims the APK — arm64
covers essentially every modern Android device, and dropping the other
ABIs halves the packaging work.
One warning: panic = "abort" means a Rust panic becomes a process
abort instead of an unwinding exception. That is fine for a mobile
app's release build, but do not enable it globally in debug, and make
sure your FFI boundary is written to never panic — FRB-generated code
already wraps calls defensively, but your own #[frb] functions should
return Results rather than panicking.
Drive Development in Release Mode — Kill the FFI Latency
The second biggest quality-of-life win has nothing to do with
compilation. Flutter+Rust debug builds are slow at runtime, not just
at build time. An unoptimized Rust core with opt-level = 0 and no LTO
can add tens of milliseconds of latency per FFI call, and when your
core is an AI agent streaming token deltas, that latency is felt on
every single interaction.
PupPal's daily development loop therefore runs the app in release mode:
flutter run --release
The first build is slower — release does LTO and codegen-units = 1,
which are inherently slower to produce — but every subsequent build
uses the incremental cache, and every runtime FFI call drops from
roughly a hundred milliseconds to a few. Hot reload still works for
Dart changes, which is what you edit most of the time. When you
genuinely need Rust breakpoints, keep one debug configuration around
for that session, but make release the default driving mode. The
discipline is simple: you pay compile time once up front so that every
runtime interaction for the rest of the day is fast.
There is a second build-time lever hiding here. The upstream
pi_agent_rust crate uses vergen-gix in its build.rs, which reads
git history on every build. In a Flutter context that can force the
build script to re-run when it does not need to. Setting
VERGEN_GIT_DISABLE=1 in the environment skips the git queries
entirely:
set VERGEN_GIT_DISABLE=1
It is a small thing, but build scripts that re-run on every build
invalidate downstream incremental compilation, and this removes one
more source of spurious rebuilds. (The permanent fix — a crate split
that removes the dependency on the CLI crate's build.rs — is Step 4
below.)
Step 3: Cache Everything — sccache and cargo check
Even with a shared target dir and tuned profiles, there are two more
caching layers worth adding. The first is sccache, a
compiler-cache daemon that works at the .rlib level:
cargo install sccache
set RUSTC_WRAPPER=sccache
With RUSTC_WRAPPER set, every rustc invocation goes through
sccache, which stores compiled artifacts keyed by inputs and reuses
them across projects and across CI runs. It is most valuable when you
have multiple workspaces sharing dependencies — exactly the PupPal
situation, where rust_puppal_api, the upstream crate, and any other
Rust tooling on the machine all pull from overlapping dependency sets.
The first build after enabling it is unchanged; every subsequent build
of a shared dependency anywhere on the machine is a cache hit.
The second habit is cheaper and often forgotten: use cargo check,
not cargo build, while iterating on Rust logic.
cd puppal/rust
cargo check
cargo check runs the type checker and borrow checker but skips code
generation, so it is typically several times faster than a full build.
For the edit-compile-fix loop — where you are correcting type errors,
not measuring performance — it is the right tool. You can even wire it
into your editor so every save runs cargo check on the bridge crate
in under a second. Only run full cargo build when you actually need
the artifact.
Step 4: Split the Giant Crate — Core and Shell
Everything so far is incremental: shared cache, tuned profiles, release driving, sccache. They make a bad situation tolerable. The structural fix that makes it good is splitting the upstream crate so the Flutter bridge never sees the CLI/TUI/example/bench machinery at all.
PupPal's plan — and the reason the earlier steps exist is to buy time
for it — is to split pi_agent_rust into a core crate and a shell:
puppal/packages/
├── pi_agent_rust/ # shell: CLI, TUI, main.rs, examples, benches
│ └── depends on pi_agent_rust_core
├── pi_agent_rust_core/ # new: only the stable SDK API Flutter needs
│ └── depends on nothing heavy
└── ...
puppal/rust/ # the FRB bridge crate
└── depends on pi_agent_rust_core
The mechanics are mechanical:
- Create
pi_agent_rust_core/with a minimalCargo.tomlthat carries only the dependencies thesdkmodule actually uses. Audit the upstreamCargo.tomland leave outclap,crossterm,bubbletea,glamour,arboard,wasmtime, theswc_*family,ast-grep-*,image,tikv-jemallocator, and every[[bin]],[[example]], and[[bench]]entry. - Migrate the SDK code as-is. Copy the
sdkmodule (and the handful of bridge functions that call into it) file by file intopi_agent_rust_core/src/, preserving signatures, fields, and error types. Where a dependency is missing, mark it with a// FIXMEplaceholder rather than silently changing behavior. - Turn the shell into a re-export. In
pi_agent_rust, addpi_agent_rust_core = { path = "../pi_agent_rust_core" }andpub use pi_agent_rust_core::*;so the CLI and TUI still compile against the same API. - Point the bridge at core. Change
rust_puppal_api's dependency frompi_agent_rusttopi_agent_rust_coreand fix theuse pi::…paths inpi_agent.
The payoff is not that core becomes small — it still contains the real
agent logic. The payoff is that editing business logic in core no
longer compiles swc or the TUI stack, and editing the CLI or
examples no longer triggers a rebuild of the bridge at all. The
thirty-minute worst case becomes a few-minute worst case, and the
common case — touching one function in the agent loop — drops to
seconds because the heavy dependencies are simply not in the graph
anymore.
The upstream crate even hints at this design itself: its lib.rs
documents sdk as the only stable library-facing surface, and says
external consumers should treat everything else as unstable. The core
crate is just that documented contract made physical.
Step 5: The Build Habits That Actually Protect You
Optimizations decay if your habits undo them. PupPal codified a small set of rules that are worth stealing wholesale:
| Don't | Instead |
|-------|---------|
| flutter clean at the first sign of trouble | flutter pub get; clean only Dart caches |
| Delete target/ by hand | Point CARGO_TARGET_DIR outside the project |
| Float the Rust toolchain | Pin it in rust-toolchain.toml |
| Build from different directories | Keep CARGO_TARGET_DIR fixed everywhere |
The toolchain pin deserves emphasis. PupPal's vendored pi_agent_rust
pins nightly-2026-07-05 in rust-toolchain.toml, with a comment
explaining exactly why: the codebase needs nightly for
feature(portable_simd) in its SQLite dependencies, and an unpinned
nightly lets clippy drift across toolchains — a newer nightly once
re-flagged nineteen clippy::unused_async sites that already carried
crate-level #![allow] attributes, turning CI red without any code
change. Pinning a known-good date keeps lints, builds, and cache
fingerprints stable. If you bump the pin, do it deliberately and re-run
clippy.
Also worth saying: the cache disciplines compound. A shared target dir
means your editor's cargo check warms the cache your build uses. Your
CI's cargo check warms the cache your local build uses (if CI shares
CARGO_TARGET_DIR or sccache). The more you use the cache, the more
things are cache hits, and the less often you ever see a cold build.
The Results and the Order of Attack
Measured on PupPal's actual hardware, the combined effect of these steps:
- Release
.sosize: 22 MB → 6–9 MB.opt-level = "z", LTO,strip, plus arm64-only packaging. flutter cleanno longer causes half-hour rebuilds. Cargo artifacts live outsidebuild/; the cache survives cleans.- Runtime FFI latency: ~100 ms → single-digit milliseconds per call in daily release-mode driving.
- The structural ceiling drops once the crate split lands: the heavyweight dependency graph leaves the Flutter build entirely.
If you are starting from scratch on your own project, the order is: set
CARGO_TARGET_DIR today, delete the old target once, and drive in
release mode; tune profiles tomorrow; do the crate split this week;
keep sccache, cargo check, and the habit table from then on. Steps
one and two alone remove the "waiting half an hour after a clean" pain
that makes people give up on Rust-in-Flutter entirely. The rest is
about making the happy path fast, not just the disaster path
survivable.
PupPal's whole point is that serious engineering can live inside a pet care app. The AI pipeline that analyzes every photo check-in, the C7 anomaly engine that watches each result, and the care codes that let someone else watch over the pet all run on top of this exact bridge. The same story applies to the developer experience: if you want to understand what makes PupPal different, it is the willingness to carry real complexity — and to engineer the build system until that complexity stops hurting. A Flutter+Rust bridge is only a good idea if you can actually develop against it. With a shared target directory, tuned profiles, and a core/shell split, you can.