
On-device AI removes your API bill and replaces it with three costs almost nobody prices in: download size, device memory, and the users you lose before they ever see the feature. Those costs are real, they are measurable, and most teams discover them after the architecture is locked.
Almost every article about on-device AI leads with what it saves. This one leads with what it costs, because that is the number that decides whether the architecture survives contact with your install base. It is also the first question worth asking before you brief Flutter developers who have shipped on-device inference.
Three things, in rough order of how much trouble each causes.
The runtime is a rounding error. The weights and the memory are the whole conversation.
Google has published first-party data on this, and it is the clearest number in mobile engineering economics.
In Google Play internal analysis, smaller download sizes correlated with higher install conversion rates. For apps under 100MB, every additional 6MB of download size correlated with a 1% decrease in install conversion. An app at around 10MB had a download completion rate roughly 30% higher than an app at around 100MB. In emerging markets the effect is sharper: removing 10MB correlated with an install conversion increase of about 2.5%, and around 70% of people in those markets said they considered app size before installing, driven by data cost and storage limits.
Google states the mechanism plainly in its own guidance: increasing your app size can negatively impact install success and increase uninstalls. Play Console surfaces this directly, reporting your download size against peers, how many active devices have under 2GB of free storage, and your uninstall ratio on those low-storage devices.
There is also a hard visibility threshold. Play Console documentation states that if your app is above 200MB, users installing over a mobile data connection see a dialog warning them about the size. It does not block the install. It does give every user a moment to reconsider.
Here is the part that most teams get wrong, and it is the reason the Google size data cannot be applied directly.
Nobody ships a 3GB model inside the app bundle. The standard pattern, and the one the flutter_gemma documentation and every serious community implementation use, is to download the weights on first launch and cache them on device. Your store listing stays small. Your install conversion rate is untouched.
The loss does not disappear. It moves.
Instead of losing users at the store page, you lose them at first launch, in front of a progress bar, over a connection you do not control, waiting for a file larger than most mobile games. The user has already installed. They have already given you their attention. Then you ask them to wait several minutes before the app does anything useful.
That is a worse place to lose someone, for two reasons. First, store conversion is a number your marketing team watches obsessively and can A/B test. First-launch activation drop-off usually sits in a blind spot until retention numbers come in. Second, a user who abandons during the model download has installed, failed to activate, and will very likely uninstall, which is the exact signal Play Console tracks and Google says affects your metrics.
Published activation benchmarks put store-page-to-install conversion at roughly 25% on iOS and 27% on Google Play. Install-to-activation above 25% is considered median. A multi-gigabyte blocking download at launch attacks the second number, and the second number is the one that compounds into retention.
More than your emulator has, which is the first thing most teams learn.
Community implementations running Gemma-class models through flutter_gemma consistently specify a physical Android device with at least 4GB of RAM, and note that emulators will not run the model at all. On iOS, published hackathon and production projects target iOS 16 and above on physical hardware. These are community figures rather than an official support matrix, so treat them as a floor to validate rather than a specification to quote to a client.
Two consequences follow, and both are architectural rather than cosmetic.
Your development loop gets slower. Every inference change needs a physical device. Your CI cannot test the AI path on a standard emulator runner. Budget for device labs and for the engineering time that simulator-based iteration used to save you.
Your addressable device base shrinks. Android fragmentation means a meaningful share of active devices in many markets sit below the memory line. Those users do not get a degraded AI experience. They get a crash, or a feature that silently fails to initialize, unless you build the fallback path deliberately.
This is the same constraint that shapes cross-platform builds across a fragmented device base, only sharper.
Model the loss before you build, not after. The shape is predictable.
| Cost | Where it lands | What to measure |
| Runtime in the binary | Store conversion | Download size delta, install conversion rate |
| Model download on first run | Activation | First-session completion rate, time to first useful action |
| Storage held after install | Retention | Uninstall ratio on devices with under 2GB free |
| RAM floor | Addressable base | Share of active devices below 4GB RAM |
| Physical-device testing | Engineering velocity | Cycle time on AI-path changes |
The pattern that matters: on-device AI does not cost you one thing once. It applies a small tax at four separate points in the funnel, and those taxes multiply rather than add.
In four situations the arithmetic is clearly favorable.
In the opposite case, where the feature is occasional, the task is open-ended, and your users are on mixed hardware, an API call is almost always the cheaper total-cost answer, even with the invoice attached.
Five moves, roughly in order of impact.
Pick the smallest model that passes your evaluation, not the best one available. Teams reach for the largest model that fits, then optimize downward. Reverse that. Define the evaluation first, then step up through model sizes until it passes, and stop there. The gap between 284MB and 4GB is the gap between a feature that ships and one that quietly gets cut.
Make the model download optional and deferred. Ship the app fully functional without AI. Let the user opt in to the download when they first reach a feature that needs it, with the size stated up front. This protects activation and converts a forced cost into an informed choice.
Build the fallback path on day one. Every device below your memory floor, and every user who declines the download, needs a working route through your product. If that route is an API call, you now have a hybrid architecture, which is usually the correct answer anyway.
Split responsibilities rather than choosing a side. The strongest published pattern runs classification, tagging, and summarization on device while sync, storage, and heavy generation stay in the cloud. The private data never leaves. The expensive general reasoning stays where it is cheap to run.
Fold this decision into the architecture phase when planning a cross-platform build, not after the model is chosen.
Instrument the four funnel points before launch. You cannot optimize what Play Console is measuring and you are not. Dynamic delivery work at Halodoc, a healthcare platform, cut Play Store install size by 40% and reported an 11% higher install conversion rate with uninstalls down 52%. Size work pays, but only if someone is watching the number.
On-device AI is not free. It trades a variable cost you can see on an invoice for a set of fixed costs spread across your install funnel, your device support matrix, and your engineering cycle time. For compliance-bound, high-frequency, or offline-first products, that trade is strongly worth making. For an occasional convenience feature on a consumer app with a broad device base, it usually is not.
The teams that get this right size the model to the task and make the download a choice. The teams that struggle pick the most capable model available, bundle it, and find out what it cost them a quarter later. Either way, the engineering time is the larger number, so it is worth knowing what a Flutter team actually costs before you commit to the harder architecture.
The runtime adds a modest amount to your binary. The model weights are the real cost and typically are not bundled. Published footprints for Flutter-compatible on-device models range from around 284MB for a small function-calling model up to 3GB to 6GB for a multimodal model. Most teams download the weights on first launch rather than shipping them in the app bundle.
Yes. Google Play analysis found that for apps under 100MB, every additional 6MB of download size correlated with roughly a 1% drop in install conversion rate, with a stronger effect in emerging markets. Google also states that increasing app size can negatively impact install success and increase uninstalls.
Community implementations running Gemma-class models in Flutter consistently specify a physical Android device with 4GB of RAM or more and note that standard emulators cannot run the model. Validate the floor against your own target devices rather than treating any published figure as a support matrix.
It depends on call frequency. On-device removes per-call cost entirely but adds fixed costs in download size, activation drop-off, device support, and testing time. High-frequency features usually favor on-device. Occasional features usually favor an API.
Yes. The flutter_gemma package family supports Android, iOS, web, and desktop, with inference engines available as opt-in modules. The cross-platform layer is not the hard part. Device capability and model size are.
Ship the app fully functional without AI, then let users opt in to the model download at the point they first need the feature, with the size stated clearly. This keeps your store conversion and first-session activation intact and turns a forced wait into an informed choice.
If you have already shipped on-device inference and the numbers are not landing where you expected, a Flutter app audit will tell you which of the four funnel points is actually leaking.
Let’s begin a collaborative journey together where we craft Flutter applications that set benchmarks for you.