Cloud Service for Connected Cars
Cars running Android Automotive OS (AAOS) retrieve information from and store information to the cloud. Features include software updates (OTA) and LLM services possibly provided by third parties such as Anthropic or Google Gemini.
High-level architecture
1. Requirements / scale
Scope
Assumptions: about 10 million vehicles worldwide, a 10–15 year vehicle lifetime, cellular links that are intermittent and costly per GB, and deployment across several regions (US, EU, China) with different data laws.
Sizing estimate - data pipeline
Suppose 30% of the fleet is driving at peak and each car sends one batched message every 10 seconds. That's 3M × 0.1 = about 300K messages/sec, and at around 1 KB each, roughly 300 MB/s of ingest. For OTA, a 2 GB full image pushed to 10M cars is 20 PB of egress. Delta updates can cut that by roughly 10x. Both numbers push you toward batching, compression, CDNs, and staged rollouts.
non-functional requirements
- High availability for safety-adjacent features like remote unlock and eCall.
- Strong security — a compromised fleet is a physical safety problem, not just a data breach.
- Regulatory compliance — UNECE R155 (cybersecurity management), R156 (software update management), ISO/SAE 21434, and GDPR.
2. Services
| Service | Responsibility | Key design choice |
|---|---|---|
| Identity & provisioning | Gives each car a unique X.509 cert, provisioned at the factory, with the key held in an HSM/TEE. Handles rotation and revocation. | mTLS everywhere; the car's identity is the root of all authorization |
| Vehicle registry / digital twin | Stores each car's VIN, hardware config, software versions per ECU, and its last-known state, kept as desired state vs. reported state. | Apps read the twin, not the car, so they work while the car is offline |
| Command service | Remote lock/unlock, climate, horn and lights, "find my car" | Async with ack, idempotency keys, TTL expiry (a 3-hour-old "unlock" must not execute), and user authorization checks |
| Telemetry pipeline | MQTT → Kafka → stream processing → time-series DB plus data lake | Priority lanes: crash/DTC alerts go immediately; bulk trip data goes batched or Wi-Fi only |
| OTA service | Builds, signs, and distributes packages; manages campaigns and tracks the update state machine per car | See deep dive below |
| LLM gateway | Brokers every AI request from the car to Anthropic, Gemini, or an on-device model | See deep dive below |
| Config & feature flags | Remote config, gradual feature enablement, A/B experiments, regional feature gating | Signed configs, cached on the car with last-known-good fallback |
| User/account & consent | Owner and driver profiles, shared cars, data-sharing consent, subscriptions | Consent state must gate what telemetry is even collected |
| Notification | Pushes to phone apps (charging done, alarm triggered) | Decoupled via events from the twin/telemetry |
| Content services | Maps/traffic, media, POI, app store for AAOS | Mostly third-party, fronted by your gateway |
| Observability & VSOC | Fleet health dashboards, plus a Vehicle Security Operations Center for intrusion detection, which R155 effectively requires | Anomaly detection on connection patterns and signing failures |
On the car side
AAOS apps shouldn't each talk to the cloud independently. A single privileged vehicle agent (a system service) should own the connection, certificates, offline queue, and retries, and expose a local API to apps. Vehicle signals come through the Car API / VHAL. Update orchestration happens in a separate component that handles both the Android A/B system update (update_engine) and flashing the other ECUs, which usually goes through the TCU or a central gateway ECU.
3. Design principles
Offline-first and intermittent connectivity
Cars enter tunnels and parking garages and drive through rural dead zones. Every interaction should be asynchronous, store-and-forward, and idempotent. The digital twin with desired and reported state is the core pattern: the cloud writes the desired state ("doors locked"), and the car reconciles and reports back whenever it reconnects. MQTT with persistent sessions and QoS 1 is a natural fit, since one long-lived connection is cheaper than repeated TLS handshakes.
Security as a chain of trust
Cover every link: hardware-backed keys, secure boot, mTLS, signed firmware, and signed configs. For OTA, cite Uptane, the automotive extension of TUF. It separates the director and image repository roles, so compromising one server key can't push malicious firmware. Apply least privilege: a compromised infotainment app should never be able to reach powertrain ECUs.
Safety boundaries, especially for the LLM
The LLM must never control safety-critical functions like braking, steering, or ADAS. Tool calls are allowlisted and routed through a policy layer. Some actions require confirmation ("open the trunk?"), and some are blocked while driving. Driver distraction rules also shape the UX: short spoken answers, no long text on screen while moving.
Fleet-scale thundering herd
Millions of cars start around 8am, and a server-side outage causes mass reconnects when it ends. Use jittered exponential backoff, connection rate limiting at the edge, and staggered OTA polling.
Never brick a car
A/B partitions with automatic rollback on failed boot are the minimum. Check preconditions before installing: parked, sufficient battery, and user consent for anything that disables the car during install. Roll out in stages — internal fleet → 1% → 10% → 100% — with automatic halt on elevated failure or crash metrics.
Cost awareness
Cellular data is a real line item for OEMs. Compress and batch uploads, use delta updates, route large downloads over Wi-Fi when possible, and let the backend dial telemetry sampling rates up or down remotely.
Data residency and privacy
Location data is highly sensitive. Deploy regional stacks (EU data stays in the EU; China typically requires a fully separate deployment), pseudonymize at ingest, enforce consent, and set retention policies.
Longevity and versioning
A 2026 car will still be calling your APIs in 2040. Version APIs explicitly, keep backward compatibility for a long time, and have the twin record exact software versions so the backend knows what each car supports.
4. OTA updates
The car reports each state transition: downloaded → verified → installing → success / failed / rolled back. The campaign dashboard aggregates these states, and automated guards pause the rollout if the failure rate crosses a threshold. Targeting matters because vehicle configurations vary enormously.
Trade-offs
- Pull vs. push — pull is simpler and scales better; push is more immediate. Typical answer: pull with a push "hint."
- Delta updates save bandwidth but require the backend to know exactly which base version each car has.
5. LLM services
The car should never call Anthropic or Google directly with provider API keys — keys embedded in a car can be extracted. All requests go through your LLM gateway, which handles:
- Provider abstraction and routing — a common internal API routes by task, region, cost, and latency, with failover if one provider is degraded. It also lets you switch vendors without an OTA.
- Context assembly — enriches the prompt server-side with vehicle context: the car's manual, current state from the twin, location, and user preferences. The car sends less data.
- Tool calling with policy enforcement — the model may propose "set cabin temp to 21°C." The gateway checks that against an allowlist and the vehicle's driving state, then sends it through the normal command service. No separate side channel into the car.
- Privacy — PII redaction, respecting consent, and data-retention agreements with providers (e.g., zero data retention where available).
- Quotas, metering, and cost control — per-vehicle and per-subscription limits, semantic caching for common questions ("how do I pair my phone"), and usage metering if AI is a paid feature.
- Streaming — voice UX needs low time-to-first-token, so stream responses and start TTS on the first sentence.
On-device fallback: a small local model on the car handles offline scenarios and simple intents ("call home," "turn up the volume") with near-zero latency. The cloud handles complex reasoning. This hybrid split addresses both connectivity and cost.
6. Storage choices
| Data | Store |
|---|---|
| Registry and twin | Strongly consistent key-value/document store keyed by VIN (DynamoDB, Cosmos DB, Cassandra) |
| Telemetry | Kafka as buffer → time-series DB for hot data → object storage data lake for history and ML |
| OTA artifacts | Object storage behind a CDN |
| Campaigns and accounts | Relational database (transactional, modest volume) |
7. What's special about cars
intermittent connectivity, safety, security, long device lifetimes, and cost per byte.
- The LLM never touches safety-critical controls
- A failed update must never brick the car
Final Check-list
- Requirements and estimates
- High-level diagram
- Two deep dives (OTA and the LLM gateway are the most interesting here)
- Failure modes and trade-offs