News and views from members of the Java team at Oracle
Starting with JDK 28, the Foreign Function and Memory (FFM) API can serve small allocations from reusable native-memory pools when applications use Arena.ofConfined().
One attractive property is that this requires no source code changes, while substantially reducing the cost of common FFM usage patterns.
Confined arenas are often used as temporary scratch space around native calls as shown in this short example:
try (Arena arena = Arena.ofConfined()) {
MemorySegment result = arena.allocate(ValueLayout.JAVA_INT);
nativeFunction.invokeExact(result); // Writes a result in the provided segment
return result.get(ValueLayout.JAVA_INT, 0);
}
Snippet 1. A common confined-arena pattern.
This pattern provides deterministic lifetime management and strong thread-confinement guarantees. When the arena closes, its segments immediately become inaccessible. The above pattern is very common when handling errno in system calls, for example.
Before JDK 28, even a tiny allocation could involve the regular native allocator and the registration of an associated cleanup action. For a native call that, for example, only needs a pointer, long, int, or short string, this bookkeeping can cost more than actually accessing the allocated memory itself.
Source and runtime analyses performed during the development of the feature confirmed that this pattern is common. In one instrumented test-suite run, some confined arenas allocated no native memory, while more than 99.99% of confined arenas that did allocate memory used less than 64 bytes. Although a test-suite workload is not representative of every application, the distribution strongly favored a small, lazy pool.
Each platform thread can lazily maintain a small cache of native-memory pools. By default, the cache can retain four pools of 64 bytes each.
A confined arena does not acquire a pool when it is created. Instead, acquisition is deferred until the arena receives its first allocation that can benefit from pooling. This is important because arenas that never allocate memory do not cause a native pool to be created. Once acquired, the pool belongs exclusively to that arena. After that, suitable allocations are carved sequentially from the pool. Here is an illustration of carving:
64-byte pool
0 8 16 20 63
+----------+----------+-----+-------------------- ... --+
| pointer | long | int | remaining capacity ... |
+----------+----------+-----+-------------------- ... --+
Figure 1. Carving out small allocations.
If an allocation does not fit, or requires an alignment that the pool cannot guarantee, it follows the regular allocation path. This makes pooling an opportunistic optimization: large and unusually aligned allocations continue to work as before.
When the arena closes, the used portion of its pool is cleared and returned to the thread cache. If the cache is already full, the pool is released to the native allocator instead. Each arena that acquires a pool owns it exclusively, so nested confined arenas (i.e., where a try-with-resource block with a first arena contains yet another try-with-resource with a second arena) remain isolated. If no cached pool is available, the arena attempts to allocate a local pool of its own. On close, that pool is either added to the cache or released.
Creating a native-memory cache in every virtual thread would undermine their small footprint. Instead, a virtual thread temporarily acquires a pool from its current carrier thread.
The virtual thread is pinned only during the short cache handoff. The pool is then detached from the carrier cache and owned by the arena. Consequently, the virtual thread can migrate safely while the arena remains open. When the arena closes, the pool is returned to the cache of the (possibly different) carrier on which the virtual thread is then mounted.
This design avoids adding a native-memory cache to potentially millions of virtual threads while preserving exclusive pool ownership.
Pooling does not change the lifetime or accessibility rules of the Arena API.
After an arena closes, this is still true:
Its scope is no longer alive.
Its memory segments cannot be accessed.
Before a pool becomes available for reuse, the portion used by the previous arena is zeroed out. Newly allocated pools are also zero-initialized. This preserves the API requirement that native segments returned by an arena contain zeroes and prevents data written by one arena from being exposed through a later arena.
The regular cleanup actions run before the pool is cleared and released. This ordering matters because a cleanup action may legitimately access the native region through a cleanup segment.
Available cached pools are deterministically released when their platform or carrier thread terminates. Pools belonging to open arenas are detached from the cache and are therefore not accidentally freed by carrier termination or virtual-thread migration.
Although segment lifetime is unchanged, the lifetime of the underlying native allocation can be different. Closing a pooled arena can return its memory to the JDK cache instead of immediately invoking the native allocator. Applications and diagnostic tools must not rely on a one-to-one correspondence between arena allocation and native malloc or free calls.
Development benchmarks showed large reductions in latency for tiny confined allocations. The following table reports nominal speedups for allocations of 5 and 20 bytes:
| Platform | 5-byte allocation | 20-byte allocation |
|---|---|---|
| Linux AArch64 | 12.3x | 8.6x |
| Linux x64 | 6.8x | 7.8x |
| macOS AArch64 | 18.6x | 16.8x |
| Windows x64 | 16.8x | 18.9x |
Table 1. Nominal speedups across the tested platforms.
Each entry is the baseline JMH average-time latency divided by the pooled latency; higher is better. The benchmark configuration and raw results are available in OpenJDK PR #31365.
A separate benchmark run on an Apple M4 running macOS produced:
Benchmark (size) Mode Cnt Score Error Units
AllocTest.alloc_confined 5 avgt 30 1.052 ± 0.027 ns/op
AllocTest.alloc_confined 20 avgt 30 1.138 ± 0.023 ns/op
AllocTest.alloc_confined 100 avgt 30 19.044 ± 0.322 ns/op
AllocTest.alloc_confined 500 avgt 30 26.383 ± 0.714 ns/op
AllocTest.alloc_confined 2000 avgt 30 31.292 ± 0.253 ns/op
AllocTest.alloc_confined 8000 avgt 30 81.119 ± 0.957 ns/op
AllocTest.alloc_confined_no_pool 5 avgt 30 15.862 ± 0.226 ns/op
AllocTest.alloc_confined_no_pool 20 avgt 30 18.139 ± 0.262 ns/op
AllocTest.alloc_confined_no_pool 100 avgt 30 18.856 ± 0.281 ns/op
AllocTest.alloc_confined_no_pool 500 avgt 30 26.168 ± 0.232 ns/op
AllocTest.alloc_confined_no_pool 2000 avgt 30 31.209 ± 0.314 ns/op
AllocTest.alloc_confined_no_pool 8000 avgt 30 80.840 ± 0.804 ns/op
Listing 1. Benchmarks on an M4 Mac.
For allocations above the default 64-byte pool size, performance remained broadly around parity because those allocations use the regular path. Virtual threads incur some additional carrier-handling overhead, but still showed substantial improvements over non-pooled allocation:
Benchmark (size) Mode Cnt Score Error Units
AllocTest.OfVirtual.alloc_confined 5 avgt 30 3.222 ± 0.087 ns/op
AllocTest.OfVirtual.alloc_confined 20 avgt 30 2.737 ± 0.021 ns/op
AllocTest.OfVirtual.alloc_confined 100 avgt 30 18.868 ± 0.199 ns/op
AllocTest.OfVirtual.alloc_confined 500 avgt 30 26.117 ± 0.240 ns/op
AllocTest.OfVirtual.alloc_confined 2000 avgt 30 31.407 ± 0.508 ns/op
AllocTest.OfVirtual.alloc_confined 8000 avgt 30 80.474 ± 0.676 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 5 avgt 30 15.672 ± 0.090 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 20 avgt 30 18.109 ± 0.250 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 100 avgt 30 18.791 ± 0.163 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 500 avgt 30 25.947 ± 0.174 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 2000 avgt 30 31.206 ± 0.440 ns/op
AllocTest.OfVirtual.alloc_confined_no_pool 8000 avgt 30 79.711 ± 0.576 ns/op
Listing 2. Virtual thread benchmarks on an M4 Mac.
As always, microbenchmark speedups do not translate directly into equivalent whole-application speedups. The largest benefits are expected where arena creation, native allocation, and cleanup form a significant part of a short native operation.
This optimization is particularly relevant for:
FFM bindings that create a confined arena around each native call
Generated jextract-style wrappers
Small C values and structures
Native handles and pointer out-parameters
Short strings and byte arrays passed to native code
Latency-sensitive code making frequent, inexpensive native calls
Note that these benefits only apply to confined arenas and not the other arena types.
The most useful performance improvements are often those that do not require developers to choose a new API or rewrite working code. Pooled confined arenas preserve the deterministic lifetime, zero-initialization, and thread-confinement properties of Arena.ofConfined().
The difference is underneath: many small, short-lived allocations can now avoid repeated trips through the native allocator.
For FFM applications that use confined arenas as scoped scratch space, the existing code retains its safety properties while running considerably faster.
Here are some links if you are interested in learning more about pooled confined arenas: