The new family of Arm C1 cores marks a major shift in the mobile and ultraportable device ecosystem, replacing the familiar Cortex with a clearer focus on sustained performance and efficiency. This generation comes with the Lumex platform and with an obvious objective: to accelerate AI on the device itself without compromising battery or temperature.
Beyond the name change, the proposal combines Armv9.3-A architecture, a profound redesign of the memory subsystem, and a significant boost to matrix computing capabilities. The result is widespread performance improvements with lower power consumption, as well as a roadmap designed for smartphones, tablets, laptops, and wearables.
Architecture and new features of the Arm C1 cores

The C1 series is organized into four variants: C1-Ultra (maximum performance), C1-Premium (high performance in less area), C1-Pro (balance) and C1-Nano (maximum efficiency). Each manufacturer can combine these blocks in heterogeneous clusters to create SoCs adapted to different ranges and uses, with configurations of up to 14 cores.
Arm has tweaked both the front- and back-end, including improvements to prediction, caches, and out-of-order execution. Thanks to the new interconnect and a more efficient (data-intensive) shared cache, SLC cells), the platform offers average increases close to 15% in everyday uses, which scale to +30% on demanding loads and reach peaks of up to 45% in multicore.
Memory support evolves with L to reduce power consumption and latency, while maintaining compatibility with LPDDR5X at speeds up to 9600 MT/s. This memory base, along with the cluster redesign, reinforces sustained performance and response under thermal stress.
C1-Ultra: the performance ceiling
As a top-of-the-range core, C1-Ultra It is targeting flagship SoCs and high-demand tasks such as computational photography, large AI models, or mobile AAA games. Compared to the Cortex-X925, Arm is talking about a +25% in single thread, a figure that helps scale overall performance when combined with more cores in the cluster.
The front-end improves the bandwidth of L1 of instructions and prediction accuracy, while the back-end increases the out-of-order execution window by around 25%, reaching around 2.000 instructions simultaneously. In addition, the L1 data capacity is doubled to 128 KB and L1 read speed is accelerated by approximately 33%.
C1-Premium: high performance in less area
For premium devices that don't need the absolute maximum, C1-Premium maintains an architecture very close to Ultra but with a 35% area reductionIt is designed to balance performance and cost, facilitating more compact designs without sacrificing significant figures.
C1-Pro: Balance and Multi-Core Muscle
In the central segment, C1-Pro replaces the Cortex‑A725 with a +11% efficiency at the same consumption and with efficiency improvements that reach up to 26% less energy at the same performanceIn gaming, Arm cites profits of around + 16 % in this class of nuclei.
The keys are in a more capable front-end (refined static prediction and a Much larger BTB), and a backend with more bandwidth in L1D and lower latency in L2 when the prediction is correct. The predictor has also been tuned to speed up response in real-world scenarios.
C1-Nano: efficiency above all else
For light tasks and extreme savings, C1-Nano increases efficiency by around 26 % compared to its predecessor (keeping the area virtually intact, ~+2% over A520). Prediction and fetch stages have been decoupled to bring instructions to L1 sooner and reduce waits for failed predictions.
In addition, the vector processing, drives are shut down when the pipeline gets stuck and traffic between L3 and DRAM is reduced (around 21% on average and up to 39% under certain loads), which alleviates consumption and improves response.
C1-DSU: Flexible clusters and lower consumption
The new C1‑DSU orchestrates the connection of the cores under a shared L3 cache and bridges the gap with the rest of the SoC (RAM, GPU, etc.). Compared to previous iterations, the design reduces typical system power consumption by around one 11% and the impact of memory by ~7%, relying on modes such as L3 Quick Nap to minimize losses when not in use.
Another key piece is the integration of the SME2 accelerators as elements external to the core: in C1-Ultra and C1-Premium their presence is mandatory, while in C1-Pro and C1-Nano It is optional depending on the manufacturer's design. Any core in the cluster can access them when present, enabling very diverse combinations (e.g., 2× C1‑Ultra + 6× C1‑Pro with one or two SME2 accelerators, or more modest combinations mixing Pro and Nano).
The Lumex platform also includes a new generation of GPUs. Although the focus of this news is on the CPUs, the Mali G1 accompanied by ~20% improvements in graphics performance, doubles the throughput of ray tracing and reduces power costs per frame by around 9%, bolstering the mix for GPU-first games and AI workloads.
SME2 and the role of the CPU in AI

The big leap in AI comes with SME2 (Scalable Matrix Extension 2), which accelerates matrix multiplications, multi-predicates, and new data types (including compact precisions like 2b/4b), and coordinates with SVE2 for advanced vectorization. In aggregate numbers, Arm talks about average improvements of 3,7x with consumption declines close to one 27%.
In practical cases, the company has shown latency reductions of 4,7x in speech recognition (Whisper Base), 2,4–2,8x speedups in text to speech and large increases in token generation for LLM (e.g. Gemma 3) that are close to × 5Running on the CPU avoids transfers to other accelerators, which reduces waiting times and provides responsiveness.
For small or interactive loads, the CPU takes center stage again: with EMS2Many everyday tasks (local image enhancement, segmentation, classification, camera effects, or audio) are completed faster, with less overhead and without going through the network. When demand increases, the GPU or an external NPU can continue to take over, but the CPU is no longer a bottleneck.
Software support is also available: there is integration in Linux and Android 16, optimized toolchains and libraries (KleidiAI), and compatibility with engines such as Unity and Unreal EngineThis will make it easier for apps and games to quickly adopt these improvements as the first commercial SoCs arrive.
Platform Lumex CSS puts all the pieces together (C1 CPU, Mali G1 GPU, interconnect and memory) with production-ready designs 3 nm, hardware telemetry and Arm system compatibility with LPDDR6. This allows partners to accelerate their mobile and laptop projects with scalable clusters of up to 14 cores and on-device AI capabilities.
The Arm C1 combines sustained performance, efficiency and a real push for AI on CPUs thanks to SME2; they offer the flexibility of C1-DSU to adapt clusters to each product range and constitute a solid foundation for the next wave of mobile and portable SoCs that seek to balance power, autonomy, and AI capabilities without always depending on the cloud.