Flutter Rust Bridge: Cutting a 30-Minute Build to Seconds — PupPal

2026-08-09

Flutter Rust Bridge: Cutting a 30-Minute Build to Seconds — PupPal

Flutter gives you a beautiful cross-platform UI in hours. Rust gives you a performance-critical core you can trust. Putting them together is supposed to be the best of both worlds, and with flutter_rust_bridge it genuinely is — until the first time you change one line of business logic in the Rust crate and wait thirty minutes for the Android build to finish. That was the exact situation PupPal hit. PupPal embeds a full AI coding-agent runtime — a Rust port of the pi-mono TypeScript agent, plus its own session bridge — inside a Flutter app, and the compile-time tax of that decision nearly derailed the project.

This post is the optimization playbook we actually executed. It is grounded in a real codebase: the rust_puppal_api crate that sits at puppal/rust/ in the PupPal monorepo, the giant pi_agent_rust crate it depends on, and the cargokit-based rust_builder plugin that wires them into the Flutter build. Every technique here is one we measured: shared CARGO_TARGET_DIR, debug profile tuning, release-mode daily driving, dependency pruning, crate splitting, sccache, and the discipline of never running flutter clean again. If you maintain a Flutter app with a Rust core, this is the order in which you should attack your own build times.

Why Flutter and Rust End Up in the Same App

PupPal is a pet-care app, but its differentiator is not the UI. It is the intelligence layer: every photo a user takes of their pet goes through an analysis pipeline, an AI agent that can reason about the photo, consult care handbooks, run tools, and stream back a natural language status update. That agent is heavy machinery. It has a session model, a tool-calling loop, provider adapters, token accounting, and a streaming event protocol. Writing that from scratch in Dart would mean reimplementing an entire agent runtime; embedding an existing one written in Rust was the only realistic path.

The architecture that resulted looks like this:

  • puppal/rust/ — the bridge crate rust_puppal_api, a thin flutter_rust_bridge (FRB) layer that exposes exactly the API the Flutter app needs. Its rust_input is crate::api, its output goes to lib/rust on the Dart side.
  • puppal/packages/pi_agent_rust/ — the upstream agent crate, a vendored copy with default-features = false so the CLI/TUI feature set never enters the mobile build.
  • puppal/rust_builder/ — a cargokit-based Flutter plugin that drives Cargo from the Flutter build for Android, iOS, Windows, macOS and Linux.

The bridge itself is small but speaks a deliberately simple protocol. In pi_bridge.rs, a PiSessionConfig struct carries the session settings from Dart — provider, model, API key, base URL, system prompt, tool-iteration budget — and a PiBridgeEvent enum streams results back as single-line JSON strings through an FRB StreamSink. That design choice (JSON strings instead of FRB custom enums) matters: it keeps the FFI surface trivially serializable and lets the Rust side evolve its event vocabulary without regenerating Dart bindings. It is a clean, minimal seam — and the compile-time monster on the other side of it was anything but minimal.

The Anatomy of a Flutter Rust Bridge Build

To understand the optimization, you have to understand what actually happens when you press build. A Flutter app with a Rust core does not compile one crate; it compiles a dependency graph, and the shape of that graph decides everything.

rust_puppal_api declares flutter_rust_bridge = "=2.12.0" and depends on pi via a path dependency:

pi = { path = "../packages/pi_agent_rust", package = "pi_agent_rust",
       default-features = false }

That one line is the whole problem in miniature. pi_agent_rust is not a small library. Its source directory contains files measured in megabytes: extensions.rs at 2.0 MB, extensions_js.rs at 1.1 MB, agent.rs and doctor.rs each over 400 KB. It depends on heavyweight crates — swc_* for JavaScript parsing, ast-grep-* for structural search, wasmtime, tikv-jemallocator, a TUI stack (crossterm, bubbletea, glamour), arboard, and more. Its own lib.rs is explicit about the deal: the crate is primarily the implementation crate for a CLI binary, and external consumers should treat everything outside the sdk module as unstable.

So every Flutter build dragged in a CLI agent's entire dependency universe, because the path dependency did not discriminate. The default-features = false flag removed the feature-gated extras, but a large fraction of the heavyweight dependencies are not feature-gated at all. A one-line change in the bridge could cascade into recompiling swc and friends. The first time we hit a full rebuild, the Android build took about thirty minutes on a fast machine. That is not a workflow; that is a daily penalty that makes you stop touching the Rust side at all.

The fixes below are ordered by effort-to-reward. Do them in this order and you will see compounding wins after the first two steps.

Step 1: Shared CARGO_TARGET_DIR — Stop flutter clean From Destroying Your Cache

The single most damaging habit in a Flutter+Rust project is flutter clean. It is the documented first resort for every Flutter build mystery, and it silently deletes the Rust target/ directory if your plugin's Gradle configuration writes Cargo artifacts under build/. Deleting the Cargo cache turns the next build into a full cold rebuild of every dependency — which, for pi_agent_rust's dependency graph, is the thirty-minute disaster.

The fix is to move the Cargo target directory out of the project tree entirely. Set a shared CARGO_TARGET_DIR environment variable pointing at a fast SSD:

# Windows
set CARGO_TARGET_DIR=D:\cargo-target-shared
# macOS / Linux
export CARGO_TARGET_DIR=/Volumes/FastSSD/cargo-target-shared

Three consequences follow. First, flutter clean cannot touch the cache anymore because the cache is not in the project. Second, the cache is reused across projects that share Rust dependencies — your bridge crate and any other Rust tooling on the machine converge on one artifact pool. Third, incremental compilation actually survives between builds, because Cargo's fingerprint database is not being deleted out from under it.

Two refinements. If you use VS Code, mirror the setting in .vscode/settings.json with "rust-analyzer.cargo.targetDir" so the language server shares the same cache. And after moving the directory, delete the old target/ inside the project once so Cargo does not get confused by two locations — then never delete it again. PupPal also went one step further in the Android build: the Gradle plugin (plugin.gradle) was changed to emit Cargo artifacts outside build/ and .gitignore was updated to match, which is what made flutter clean fundamentally incapable of nuking the Rust cache. This one change alone removed the "half-hour rebuild after every clean" failure mode permanently.

Step 2: Profile Tuning — The 10-Second Edit That Changes Everything

Cargo profiles are where you get the biggest wins for the smallest effort, and there is a subtlety that trips up almost everyone: profile settings in a dependency's own Cargo.toml are ignored when that crate is compiled as a dependency. The effective profiles are the ones in the root crate of the build — in PupPal's case, rust_puppal_api. pi_agent_rust can declare whatever it likes in [profile.dev]; when it is built as a dependency of the bridge, those declarations do not apply. You must configure profiles in the crate that Cargo treats as the workspace root.

For the debug profile, the goal is compile speed and link speed, not runtime speed:

[profile.dev]
opt-level = 0
debug = 1
codegen-units = 16
incremental = true

debug = 1 is the interesting line. Full debug info (debug = 2) is the default, and it significantly slows linking of large crates because the linker has to process enormous symbol tables. debug = 1 keeps only the line-number tables — which is all a debugger actually needs for stepping and breakpoints — while cutting link time dramatically. codegen-units = 16 maximizes parallel codegen, and incremental = true makes sure the incremental compilation cache is active.

For the release profile, the priorities flip: you are shipping this code to phones, so size and runtime performance dominate:

[profile.release]
opt-level = "z"
lto = true
codegen-units = 1
strip = true
panic = "abort"
debug = false
debug-assertions = false
overflow-checks = false

PupPal's release profile mirrors what pi_agent_rust itself ships: opt-level = "z" optimizes for binary size, lto = true with codegen-units = 1 enables whole-program link-time optimization, and strip = true removes symbol tables. The measured effect on the Android shared library was a drop from about 22 MB to the 6–9 MB range. On top of that, the app's Gradle config was changed to build release only for arm64-v8a, which further trims the APK — arm64 covers essentially every modern Android device, and dropping the other ABIs halves the packaging work.

One warning: panic = "abort" means a Rust panic becomes a process abort instead of an unwinding exception. That is fine for a mobile app's release build, but do not enable it globally in debug, and make sure your FFI boundary is written to never panic — FRB-generated code already wraps calls defensively, but your own #[frb] functions should return Results rather than panicking.

Drive Development in Release Mode — Kill the FFI Latency

The second biggest quality-of-life win has nothing to do with compilation. Flutter+Rust debug builds are slow at runtime, not just at build time. An unoptimized Rust core with opt-level = 0 and no LTO can add tens of milliseconds of latency per FFI call, and when your core is an AI agent streaming token deltas, that latency is felt on every single interaction.

PupPal's daily development loop therefore runs the app in release mode:

flutter run --release

The first build is slower — release does LTO and codegen-units = 1, which are inherently slower to produce — but every subsequent build uses the incremental cache, and every runtime FFI call drops from roughly a hundred milliseconds to a few. Hot reload still works for Dart changes, which is what you edit most of the time. When you genuinely need Rust breakpoints, keep one debug configuration around for that session, but make release the default driving mode. The discipline is simple: you pay compile time once up front so that every runtime interaction for the rest of the day is fast.

There is a second build-time lever hiding here. The upstream pi_agent_rust crate uses vergen-gix in its build.rs, which reads git history on every build. In a Flutter context that can force the build script to re-run when it does not need to. Setting VERGEN_GIT_DISABLE=1 in the environment skips the git queries entirely:

set VERGEN_GIT_DISABLE=1

It is a small thing, but build scripts that re-run on every build invalidate downstream incremental compilation, and this removes one more source of spurious rebuilds. (The permanent fix — a crate split that removes the dependency on the CLI crate's build.rs — is Step 4 below.)

Step 3: Cache Everything — sccache and cargo check

Even with a shared target dir and tuned profiles, there are two more caching layers worth adding. The first is sccache, a compiler-cache daemon that works at the .rlib level:

cargo install sccache
set RUSTC_WRAPPER=sccache

With RUSTC_WRAPPER set, every rustc invocation goes through sccache, which stores compiled artifacts keyed by inputs and reuses them across projects and across CI runs. It is most valuable when you have multiple workspaces sharing dependencies — exactly the PupPal situation, where rust_puppal_api, the upstream crate, and any other Rust tooling on the machine all pull from overlapping dependency sets. The first build after enabling it is unchanged; every subsequent build of a shared dependency anywhere on the machine is a cache hit.

The second habit is cheaper and often forgotten: use cargo check, not cargo build, while iterating on Rust logic.

cd puppal/rust
cargo check

cargo check runs the type checker and borrow checker but skips code generation, so it is typically several times faster than a full build. For the edit-compile-fix loop — where you are correcting type errors, not measuring performance — it is the right tool. You can even wire it into your editor so every save runs cargo check on the bridge crate in under a second. Only run full cargo build when you actually need the artifact.

Step 4: Split the Giant Crate — Core and Shell

Everything so far is incremental: shared cache, tuned profiles, release driving, sccache. They make a bad situation tolerable. The structural fix that makes it good is splitting the upstream crate so the Flutter bridge never sees the CLI/TUI/example/bench machinery at all.

PupPal's plan — and the reason the earlier steps exist is to buy time for it — is to split pi_agent_rust into a core crate and a shell:

puppal/packages/
├── pi_agent_rust/          # shell: CLI, TUI, main.rs, examples, benches
│   └── depends on pi_agent_rust_core
├── pi_agent_rust_core/     # new: only the stable SDK API Flutter needs
│   └── depends on nothing heavy
└── ...
puppal/rust/                # the FRB bridge crate
    └── depends on pi_agent_rust_core

The mechanics are mechanical:

  1. Create pi_agent_rust_core/ with a minimal Cargo.toml that carries only the dependencies the sdk module actually uses. Audit the upstream Cargo.toml and leave out clap, crossterm, bubbletea, glamour, arboard, wasmtime, the swc_* family, ast-grep-*, image, tikv-jemallocator, and every [[bin]], [[example]], and [[bench]] entry.
  2. Migrate the SDK code as-is. Copy the sdk module (and the handful of bridge functions that call into it) file by file into pi_agent_rust_core/src/, preserving signatures, fields, and error types. Where a dependency is missing, mark it with a // FIXME placeholder rather than silently changing behavior.
  3. Turn the shell into a re-export. In pi_agent_rust, add pi_agent_rust_core = { path = "../pi_agent_rust_core" } and pub use pi_agent_rust_core::*; so the CLI and TUI still compile against the same API.
  4. Point the bridge at core. Change rust_puppal_api's dependency from pi_agent_rust to pi_agent_rust_core and fix the use pi::… paths in pi_agent.

The payoff is not that core becomes small — it still contains the real agent logic. The payoff is that editing business logic in core no longer compiles swc or the TUI stack, and editing the CLI or examples no longer triggers a rebuild of the bridge at all. The thirty-minute worst case becomes a few-minute worst case, and the common case — touching one function in the agent loop — drops to seconds because the heavy dependencies are simply not in the graph anymore.

The upstream crate even hints at this design itself: its lib.rs documents sdk as the only stable library-facing surface, and says external consumers should treat everything else as unstable. The core crate is just that documented contract made physical.

Step 5: The Build Habits That Actually Protect You

Optimizations decay if your habits undo them. PupPal codified a small set of rules that are worth stealing wholesale:

| Don't | Instead | |-------|---------| | flutter clean at the first sign of trouble | flutter pub get; clean only Dart caches | | Delete target/ by hand | Point CARGO_TARGET_DIR outside the project | | Float the Rust toolchain | Pin it in rust-toolchain.toml | | Build from different directories | Keep CARGO_TARGET_DIR fixed everywhere |

The toolchain pin deserves emphasis. PupPal's vendored pi_agent_rust pins nightly-2026-07-05 in rust-toolchain.toml, with a comment explaining exactly why: the codebase needs nightly for feature(portable_simd) in its SQLite dependencies, and an unpinned nightly lets clippy drift across toolchains — a newer nightly once re-flagged nineteen clippy::unused_async sites that already carried crate-level #![allow] attributes, turning CI red without any code change. Pinning a known-good date keeps lints, builds, and cache fingerprints stable. If you bump the pin, do it deliberately and re-run clippy.

Also worth saying: the cache disciplines compound. A shared target dir means your editor's cargo check warms the cache your build uses. Your CI's cargo check warms the cache your local build uses (if CI shares CARGO_TARGET_DIR or sccache). The more you use the cache, the more things are cache hits, and the less often you ever see a cold build.

The Results and the Order of Attack

Measured on PupPal's actual hardware, the combined effect of these steps:

  • Release .so size: 22 MB → 6–9 MB. opt-level = "z", LTO, strip, plus arm64-only packaging.
  • flutter clean no longer causes half-hour rebuilds. Cargo artifacts live outside build/; the cache survives cleans.
  • Runtime FFI latency: ~100 ms → single-digit milliseconds per call in daily release-mode driving.
  • The structural ceiling drops once the crate split lands: the heavyweight dependency graph leaves the Flutter build entirely.

If you are starting from scratch on your own project, the order is: set CARGO_TARGET_DIR today, delete the old target once, and drive in release mode; tune profiles tomorrow; do the crate split this week; keep sccache, cargo check, and the habit table from then on. Steps one and two alone remove the "waiting half an hour after a clean" pain that makes people give up on Rust-in-Flutter entirely. The rest is about making the happy path fast, not just the disaster path survivable.

PupPal's whole point is that serious engineering can live inside a pet care app. The AI pipeline that analyzes every photo check-in, the C7 anomaly engine that watches each result, and the care codes that let someone else watch over the pet all run on top of this exact bridge. The same story applies to the developer experience: if you want to understand what makes PupPal different, it is the willingness to carry real complexity — and to engineer the build system until that complexity stops hurting. A Flutter+Rust bridge is only a good idea if you can actually develop against it. With a shared target directory, tuned profiles, and a core/shell split, you can.