Graviton5 cache availability and M4 comparison
Checked 2026-09-10 UTC on the existing c9g.4xlarge.
Reported hierarchy
Direct Linux sysfs inspection reports 49152 KiB = 48 MiB shared L3, with 64-byte lines and 16 ways, plus 2 MiB private L2 per core. The online parent CPUs 0,13–15 all report the same shared L3 domain. CPUs 1–12 remain reserved for the production Nitro enclave and do not expose cache directories in the parent. The online CPU list does not establish a separate 48 MiB budget for those four parent cores or an exclusive allocation for this EC2 instance.
AWS's hardware guide lists 48 MB per NUMA region for this size. Amazon's processor description explains that the advertised 192 MB is the total across four chiplets and that cache organization depends on VM size. Nitro documentation explains that exposed topology can describe shared last-level caches. Neither the OS report nor these measurements proves 48 MiB of exclusive usable capacity.
Capacity probe under current load
The probe ran at nice 10 pinned to parent CPU 13. The existing production enclave continued running; it was not restarted or modified. Maximum probe allocation was 128 MiB plus a small initialization permutation. There was one dependent load stream, with a randomized cycle visiting one pointer per 64-byte line.
chase.c is the first probe. chase-encoded.c XOR-encodes stored pointer values
to reduce straightforward pointer-prefetch opportunities and increases warm-up
from two to eight full traversals. This does not prove every hardware prefetcher
is defeated. Each size then gets two consecutive samples of eight million loads.
The samples share one allocation and traversal order; they are not independent
placement experiments. Allocation, shuffle and warm-up are outside timing.
Observed encoded-probe time per dependent load, including its XOR and loop:
| Working set | First sample, ns | Second sample, ns |
|---|---|---|
| 4 MiB | 5.11 | 5.11 |
| 8 MiB | 5.19 | 5.19 |
| 16 MiB | 4.78 | 4.79 |
| 24 MiB | 5.01 | 5.02 |
| 32 MiB | 10.42 | 8.92 |
| 40 MiB | 24.01 | 20.59 |
| 48 MiB | 38.27 | 31.47 |
| 64 MiB | 55.08 | 45.74 |
| 128 MiB | 72.96 | 71.25 |
All listed working sets had full anonymous huge-page coverage. The 24 MiB set stayed near 5 ns/load; 32 MiB was around 9–10 ns, and latency grew at 40–48 MiB. This supports useful effective cache capacity in the tens of MiB under current conditions. It does not identify a hard allocation boundary, partition quota, cache-hit latency, or a guaranteed capacity inside the enclave. Replacement, other workloads, prefetching and the private L2 all affect the curve.
The owner's M4 laptop
Direct sysctl results in m4-sysctl.txt show:
- Performance cluster: 16 MiB L2, shared by four performance cores.
- Efficiency cluster: 4 MiB L2, shared by six efficiency cores.
- No conventional L3 size returned by
hw.l3cachesize; this does not imply that the chip has no system-level cache.
TechInsights' M4 analysis estimates the last-level cache at 8 MB. This is an external hardware-analysis estimate, not a capacity measured on this laptop. Apple's optimization-guide link required Apple sign-in when accessed, so its contents were not independently verified here. The reported cluster L2 and estimated system-level cache must not be treated as one flat, dedicated cache pool for a CPU worker.
Artifacts
topology.json contains lscpu output. Both C sources and all original and encoded
probe results are retained. Compilation used the existing ARM toolchain image
with cc -O3 -std=c11; benchmark execution was native on the parent host. Build
containers remain stopped. No system-wide cache, page, or CPU settings changed.