For developers · Checked 1 October 2026

Apple Foundation Models: on-device vs Private Cloud Compute

Short answer: the Foundation Models framework lets your app prompt the language model behind Apple Intelligence, with guided generation of Swift types through @Generable and tool calling. The on-device model has a context window of 4,096 tokens per session, works offline and has no usage limit. From iOS 27, PrivateCloudComputeLanguageModel offers a 32K-token context and stronger reasoning, needs a network connection, gives each person a daily request limit and requires a managed entitlement. Both need a device and region that support Apple Intelligence.

On-device model vs Private Cloud Compute

From Apple’s «Adding server-side intelligence with Private Cloud Compute», read 1 October 2026.

SystemLanguageModel (on-device)PrivateCloudComputeLanguageModel
Context size4K tokens32K tokens
Works offlineYesNo
Usage limitsUnlimitedLimit per day, more with iCloud+
ReasoningNot supportedMultiple levels
Preserves privacyYesYes
AvailabilityiOS, iPadOS, macOS, visionOS 26 and later; watchOS 27iOS, macOS, watchOS, visionOS 27 and later
AccessAny app on a supported deviceManaged entitlement with eligibility requirements

Staying inside the 4K context window

Every interaction in a LanguageModelSession uses the same window: prompts, instructions, tool definitions and their input and output, generable schemas and every response. In English a token is typically three to four characters; in Chinese, Japanese, Korean and Vietnamese it is typically one character. When the window fills, the session throws contextSizeExceeded; start a new session, carrying over only the state you need. Use tokenCount(for:) and contextSize to measure, and the Foundation Models instrument in Instruments to see tokens per request. Apple recommends short imperative prompts of no more than three paragraphs, asking for less output, and maximumCount on generable arrays rather than cutting responses with maximumResponseTokens.

Using Private Cloud Compute

Switching is one line: both models conform to LanguageModel, so you pass PrivateCloudComputeLanguageModel to the session initializer and your prompts, tools and instructions carry over. There are no API keys to manage. Apple advises starting with the on-device model, evaluating your feature, and moving to PCC only if it needs more reasoning or context. Because PCC exists only on 27-series OS versions and needs a network, check availability, fall back to the on-device model on earlier versions, and retry on-device if the network request fails. To develop with PCC you must meet eligibility requirements and request the managed entitlement.

Daily limits

Each person gets a daily PCC request limit and can upgrade their iCloud+ subscription for more. Use the quota status to show whether someone is below, approaching or over their limit; when they are over it, the session throws quotaLimitReached, and resetDate says when the quota refreshes, if known. Apple asks for clear status UI rather than a dismissible alert, and the framework provides a path to system UI for upgrading. Apple’s documentation doesn’t publish the size of the daily limit.

Availability, languages and model versions

Check SystemLanguageModel.default.availability before prompting: it depends on whether the device and region support Apple Intelligence. The model is multilingual; call supportsLocale(_:) before using it and handle unsupportedLanguageOrLocale by explaining the limitation or turning the feature off. Apple updates the on-device model in routine OS updates; there are currently three versions, for 26.0 to 26.3, 26.4 and 27.0, so retest prompts on each. Related: Android AppFunctions, Google’s way to expose app features to AI agents, and how to market an AI app.

Where Censuus fits

An on-device AI feature is a reason to try your app; being found is what gets people to it. Censuus ranks apps by the visits they draw, and adding your app is free on the List my app form. Placement comes from traffic and sponsorship: apps climb on the visits they draw, and a sponsor can pay to rise higher.

Frequently asked questions

What is the context window of Apple’s on-device foundation model?

4,096 tokens per session, covering prompts, instructions, tools, schemas and responses.

How big is the Private Cloud Compute context window?

32K tokens, with stronger reasoning than the on-device model.

Is Apple’s Foundation Models framework free to use?

The on-device model has no usage limit. Private Cloud Compute gives each person a daily request limit, which they can raise by upgrading iCloud+; developers don’t manage API keys.

Which OS versions support Foundation Models?

iOS, iPadOS, macOS and visionOS 26 and later, and watchOS 27. PrivateCloudComputeLanguageModel needs iOS, macOS, watchOS or visionOS 27 or later.

Can every app use Private Cloud Compute?

No. You must meet Apple’s eligibility requirements and request a managed entitlement.

Guides for app developers

Put your app in the ranking

Censuus ranks apps by the visits they draw and by sponsorship, no bots. Listing is free; sponsorship raises placement.