Skip to main content

Cloud Service for Connected Cars

Cars running Android Automotive OS (AAOS) retrieve information from and store information to the cloud. Features include software updates (OTA) and LLM services possibly provided by third parties such as Anthropic or Google Gemini.


High-level architecture​


1. Requirements / scale​

Scope​

Assumptions: about 10 million vehicles worldwide, a 10–15 year vehicle lifetime, cellular links that are intermittent and costly per GB, and deployment across several regions (US, EU, China) with different data laws.

Sizing estimate - data pipeline​

Suppose 30% of the fleet is driving at peak and each car sends one batched message every 10 seconds. That's 3M × 0.1 = about 300K messages/sec, and at around 1 KB each, roughly 300 MB/s of ingest. For OTA, a 2 GB full image pushed to 10M cars is 20 PB of egress. Delta updates can cut that by roughly 10x. Both numbers push you toward batching, compression, CDNs, and staged rollouts.

non-functional requirements​

  • High availability for safety-adjacent features like remote unlock and eCall.
  • Strong security — a compromised fleet is a physical safety problem, not just a data breach.
  • Regulatory compliance — UNECE R155 (cybersecurity management), R156 (software update management), ISO/SAE 21434, and GDPR.

2. Services​

ServiceResponsibilityKey design choice
Identity & provisioningGives each car a unique X.509 cert, provisioned at the factory, with the key held in an HSM/TEE. Handles rotation and revocation.mTLS everywhere; the car's identity is the root of all authorization
Vehicle registry / digital twinStores each car's VIN, hardware config, software versions per ECU, and its last-known state, kept as desired state vs. reported state.Apps read the twin, not the car, so they work while the car is offline
Command serviceRemote lock/unlock, climate, horn and lights, "find my car"Async with ack, idempotency keys, TTL expiry (a 3-hour-old "unlock" must not execute), and user authorization checks
Telemetry pipelineMQTT → Kafka → stream processing → time-series DB plus data lakePriority lanes: crash/DTC alerts go immediately; bulk trip data goes batched or Wi-Fi only
OTA serviceBuilds, signs, and distributes packages; manages campaigns and tracks the update state machine per carSee deep dive below
LLM gatewayBrokers every AI request from the car to Anthropic, Gemini, or an on-device modelSee deep dive below
Config & feature flagsRemote config, gradual feature enablement, A/B experiments, regional feature gatingSigned configs, cached on the car with last-known-good fallback
User/account & consentOwner and driver profiles, shared cars, data-sharing consent, subscriptionsConsent state must gate what telemetry is even collected
NotificationPushes to phone apps (charging done, alarm triggered)Decoupled via events from the twin/telemetry
Content servicesMaps/traffic, media, POI, app store for AAOSMostly third-party, fronted by your gateway
Observability & VSOCFleet health dashboards, plus a Vehicle Security Operations Center for intrusion detection, which R155 effectively requiresAnomaly detection on connection patterns and signing failures

On the car side​

AAOS apps shouldn't each talk to the cloud independently. A single privileged vehicle agent (a system service) should own the connection, certificates, offline queue, and retries, and expose a local API to apps. Vehicle signals come through the Car API / VHAL. Update orchestration happens in a separate component that handles both the Android A/B system update (update_engine) and flashing the other ECUs, which usually goes through the TCU or a central gateway ECU.


3. Design principles​

Offline-first and intermittent connectivity​

Cars enter tunnels and parking garages and drive through rural dead zones. Every interaction should be asynchronous, store-and-forward, and idempotent. The digital twin with desired and reported state is the core pattern: the cloud writes the desired state ("doors locked"), and the car reconciles and reports back whenever it reconnects. MQTT with persistent sessions and QoS 1 is a natural fit, since one long-lived connection is cheaper than repeated TLS handshakes.

Security as a chain of trust​

Cover every link: hardware-backed keys, secure boot, mTLS, signed firmware, and signed configs. For OTA, cite Uptane, the automotive extension of TUF. It separates the director and image repository roles, so compromising one server key can't push malicious firmware. Apply least privilege: a compromised infotainment app should never be able to reach powertrain ECUs.

Safety boundaries, especially for the LLM​

The LLM must never control safety-critical functions like braking, steering, or ADAS. Tool calls are allowlisted and routed through a policy layer. Some actions require confirmation ("open the trunk?"), and some are blocked while driving. Driver distraction rules also shape the UX: short spoken answers, no long text on screen while moving.

Fleet-scale thundering herd​

Millions of cars start around 8am, and a server-side outage causes mass reconnects when it ends. Use jittered exponential backoff, connection rate limiting at the edge, and staggered OTA polling.

Never brick a car​

A/B partitions with automatic rollback on failed boot are the minimum. Check preconditions before installing: parked, sufficient battery, and user consent for anything that disables the car during install. Roll out in stages — internal fleet → 1% → 10% → 100% — with automatic halt on elevated failure or crash metrics.

Cost awareness​

Cellular data is a real line item for OEMs. Compress and batch uploads, use delta updates, route large downloads over Wi-Fi when possible, and let the backend dial telemetry sampling rates up or down remotely.

Data residency and privacy​

Location data is highly sensitive. Deploy regional stacks (EU data stays in the EU; China typically requires a fully separate deployment), pseudonymize at ingest, enforce consent, and set retention policies.

Longevity and versioning​

A 2026 car will still be calling your APIs in 2040. Version APIs explicitly, keep backward compatibility for a long time, and have the twin record exact software versions so the backend knows what each car supports.


4. OTA updates​

The car reports each state transition: downloaded → verified → installing → success / failed / rolled back. The campaign dashboard aggregates these states, and automated guards pause the rollout if the failure rate crosses a threshold. Targeting matters because vehicle configurations vary enormously.

Trade-offs

  • Pull vs. push — pull is simpler and scales better; push is more immediate. Typical answer: pull with a push "hint."
  • Delta updates save bandwidth but require the backend to know exactly which base version each car has.

5. LLM services​

The car should never call Anthropic or Google directly with provider API keys — keys embedded in a car can be extracted. All requests go through your LLM gateway, which handles:

  • Provider abstraction and routing — a common internal API routes by task, region, cost, and latency, with failover if one provider is degraded. It also lets you switch vendors without an OTA.
  • Context assembly — enriches the prompt server-side with vehicle context: the car's manual, current state from the twin, location, and user preferences. The car sends less data.
  • Tool calling with policy enforcement — the model may propose "set cabin temp to 21°C." The gateway checks that against an allowlist and the vehicle's driving state, then sends it through the normal command service. No separate side channel into the car.
  • Privacy — PII redaction, respecting consent, and data-retention agreements with providers (e.g., zero data retention where available).
  • Quotas, metering, and cost control — per-vehicle and per-subscription limits, semantic caching for common questions ("how do I pair my phone"), and usage metering if AI is a paid feature.
  • Streaming — voice UX needs low time-to-first-token, so stream responses and start TTS on the first sentence.

On-device fallback: a small local model on the car handles offline scenarios and simple intents ("call home," "turn up the volume") with near-zero latency. The cloud handles complex reasoning. This hybrid split addresses both connectivity and cost.


6. Storage choices​

DataStore
Registry and twinStrongly consistent key-value/document store keyed by VIN (DynamoDB, Cosmos DB, Cassandra)
TelemetryKafka as buffer → time-series DB for hot data → object storage data lake for history and ML
OTA artifactsObject storage behind a CDN
Campaigns and accountsRelational database (transactional, modest volume)

7. What's special about cars​

intermittent connectivity, safety, security, long device lifetimes, and cost per byte.

  • The LLM never touches safety-critical controls
  • A failed update must never brick the car

Final Check-list​

  1. Requirements and estimates
  2. High-level diagram
  3. Two deep dives (OTA and the LLM gateway are the most interesting here)
  4. Failure modes and trade-offs