
Start with a particular feature, not an AI phone label
On-device AI runs a model on your phone, tablet, or computer. Cloud AI sends a request to a model running on remote infrastructure. A hybrid system can use both. That distinction describes processing location; it does not, by itself, tell you whether an answer is accurate, whether an app keeps a history, or whether the entire experience works offline.
The useful unit of comparison is a task. Ask where a specific summary, transcription, or rewrite is processed, rather than whether a device is generally 'local AI.' Two buttons inside the same app can lead to different services. A device-level marketing description is not a complete account of every feature installed on it.
For concrete examples, Google's Gemini Nano documentation describes local model execution through Android's AICore. Apple's Intelligence privacy notice describes both on-device tasks and requests sent to Private Cloud Compute. Neither description should be applied automatically to unrelated third-party apps.
Trace the request from input to saved result
Consider a hypothetical request to shorten a note. First, the application obtains the text. Next, it prepares the material and instructions for a model. The model produces a response, and the application displays or saves it. Inference is the model-processing step; it is not the whole journey of the note.
Draw four boxes on paper: input, processing, output, and storage. Under each box, write what the feature's documentation actually says. If you only know the location of processing, leave the other boxes unanswered. This simple exercise prevents a narrow technical claim from quietly becoming a much broader privacy promise.
For instance, locally producing a summary and then saving it to a synchronized notebook are separate actions. Whether a particular notebook synchronizes, and under what settings, needs its own check. A local model does not tell you the storage policy of the app that receives its answer. Likewise, a remote model does not automatically tell you how long the provider retains a request.
What local processing means in the Android example
Google describes AICore as the Android system service through which apps can access Gemini Nano on the device. It manages the model and its updates. The documentation also distinguishes model execution from network access for tasks such as model downloads, which is routed through Private Compute Services. Gemini Nano architecture.
Google's ML Kit documentation lists local generative features including summarization, proofreading, rewriting, and image description. It says input, inference, and output processing for those APIs happen on-device. Device support is feature-specific, with different support for the Prompt API. ML Kit GenAI overview.
For a reader evaluating an app, these are useful architectural examples, not proof that the app uses that particular implementation. Look for an explicit statement from the app developer about the named feature. An Android phone supporting a local model does not establish that every assistant on that phone uses it.
Keep setup and everyday use separate in your notes. 'Can execute locally' and 'is ready to execute on this device today' are different questions. Before relying on a feature during travel, establish whether the required software and model are available and whether the exact task you need is supported.
What cloud processing means in the Apple example
Apple's privacy notice says Apple Intelligence evaluates whether a task can run on-device and can send relevant request data to Private Cloud Compute when a larger model is needed. It identifies notification previews as a local example and says Writing Tools may use server processing. Apple states that PCC does not retain the request content after processing. These are Apple's descriptions, not findings from our own audit. Apple Intelligence and Privacy.
Architecture descriptions also change. In June 2026, Apple announced an expansion of PCC to Google Cloud infrastructure using NVIDIA GPUs. The announcement described a gradual summer preview ramp toward the full protections. Do not repeat the older assumption that every PCC workload necessarily runs only in Apple's own data centers or on Apple silicon, or infer universal availability from an announcement. Apple's PCC expansion.
The practical lesson is to keep a date beside an architecture claim. A privacy notice, a launch announcement, and a feature availability page may describe different stages or scopes. If they appear inconsistent, identify the difference rather than combining their strongest promises into an imaginary universal specification.
Why hybrid processing needs a more careful question
Imagine a feature that begins with a local step and sends a more demanding request to a remote model. Calling that feature simply local loses an important part of the story; calling it simply cloud may hide useful local work. For an evaluation, ask what triggers the change, what information accompanies it, and whether the user can see or control it.
Do not assume that every hybrid product follows this hypothetical sequence. It is a way to frame questions, not a description of an untested device. The vendor's documentation should identify the actual workflow. If it does not, record the routing as unknown rather than filling the gap with an appealing explanation.
Also separate the operating system's service from an optional external integration. A statement about one provider's processing environment is not automatically a statement about another provider's chatbot. Identify the destination before comparing its controls. The word 'assistant' can describe an interface that reaches more than one underlying system.
Check offline behavior without treating it as a privacy audit
Our suggested functional check uses invented, non-sensitive material. Prepare a short sample note, complete any documented setup, and try the exact feature with connectivity disabled. Then record what happened, including the app version, device, and whether the result appeared complete or the feature declined to run. This is a proposed procedure; we have not performed it for this article.
A successful offline run can establish that the tested task worked in that condition. It cannot prove that all tasks in the app always stay local, that a previous request was not cached, or that no information is synchronized later. Avoid turning a useful functional observation into a conclusion it cannot support.
Repeat the check with another invented input if you need to rule out simply seeing an old result. Treat unexpected behavior as a question for the documentation or developer. For a trip or unreliable connection, retain a non-AI way to access essential notes rather than making an untested feature your only route to them.
Compare data handling separately from processing location
Build a small comparison record for each feature. Name the input it receives, the processing destination, any saved history, the available deletion controls, and the source for each answer. Add a separate line for analytics or optional data sharing. Use 'not established' where the documentation does not answer the question.
Apple documents an Apple Intelligence Report for examining requests sent off-device, including supported external integration requests. Its privacy notice also discusses optional Device Analytics separately from request processing. Those distinctions illustrate why one privacy slogan is not enough to describe every data flow. Apple's transparency and analytics information.
Our recommendation is to start evaluations with sample material you are comfortable disclosing. A polished interface or a familiar brand is not evidence that a particular document belongs in that feature. Before using private work material, check the relevant organizational rules and the feature's actual settings. This guide cannot establish permission for a specific employer, account, or document.
Do not infer speed or quality from the architecture label
Processing location is not a benchmark. To compare two features fairly, decide what success means for the same task: preserving the important facts in a summary, following a requested tone, or correctly describing a supplied image. Write those criteria before looking at the responses so a fluent answer does not quietly become your only standard.
Our suggested comparison includes several invented examples of different lengths and complexity, with the same instructions for both features. Record errors and omissions, not only completion time. Keep setup time distinct from response time. If you publish your results, include the device, software versions, connection conditions, and date so readers can understand the limits.
We have not measured latency, battery use, accuracy, or hardware performance for this guide. There is consequently no ranking here of local versus cloud quality. A feature that handles your specific task well is a more defensible choice than one selected solely because its architecture sounds newer.
A feature-level decision you can explain
Before depending on an AI feature, finish this sentence: for this task, on this device and software version, I know where processing happens because of this source. Then add what you know about connectivity and saved data. If the sentence contains guesses, keep those guesses visible until you can check them.
Availability and behavior should be verified for your exact feature, device, software, language, account, and region rather than inferred from a broad announcement. Record only the conditions relevant to the service you are evaluating. A global comparison is most useful when it makes those boundaries explicit instead of assuming every reader sees the same options.
On-device, cloud, and hybrid are useful starting descriptions. The stronger decision comes from tracing the actual request, distinguishing vendor claims from your own observations, and checking what happens before and after inference. That gives you a repeatable way to assess the next feature, even when the branding changes.
Sources
- Gemini Nano
Android Developers | Checked
- Overview of the ML Kit GenAI APIs
Google for Developers | Checked
- Apple Intelligence and Privacy
Apple | Checked
- Expanding Private Cloud Compute
Apple Security Research | Published | Checked
Editorial disclosure
This article was written with AI assistance in Codex. The featured image is an AI-generated conceptual illustration, not a hardware diagram. Vendor statements are attributed and were checked on September 8, 2026. The example workflows are hypothetical; no device benchmarks, network inspection, or independent security audit were performed.


