Java Cup
Inside Java

News and views from members of the Java team at Oracle

Making Arena.ofConfined() Even Cheaper in JDK 28

Starting with JDK 28, the Foreign Function and Memory (FFM) API can serve small allocations from reusable native-memory pools when applications use Arena.ofConfined().

One attractive property is that this requires no source code changes, while substantially reducing the cost of common FFM usage patterns.

Small Native Allocations Are Common

Confined arenas are often used as temporary scratch space around native calls as shown in this short example:

try (Arena arena = Arena.ofConfined()) {
    MemorySegment result = arena.allocate(ValueLayout.JAVA_INT);
    nativeFunction.invokeExact(result); // Writes a result in the provided segment
    return result.get(ValueLayout.JAVA_INT, 0);
}

Snippet 1. A common confined-arena pattern.

This pattern provides deterministic lifetime management and strong thread-confinement guarantees. When the arena closes, its segments immediately become inaccessible. The above pattern is very common when handling errno in system calls, for example.

Before JDK 28, even a tiny allocation could involve the regular native allocator and the registration of an associated cleanup action. For a native call that, for example, only needs a pointer, long, int, or short string, this bookkeeping can cost more than actually accessing the allocated memory itself.

Source and runtime analyses performed during the development of the feature confirmed that this pattern is common. In one instrumented test-suite run, some confined arenas allocated no native memory, while more than 99.99% of confined arenas that did allocate memory used less than 64 bytes. Although a test-suite workload is not representative of every application, the distribution strongly favored a small, lazy pool.

How Pooling Works

Each platform thread can lazily maintain a small cache of native-memory pools. By default, the cache can retain four pools of 64 bytes each.

A confined arena does not acquire a pool when it is created. Instead, acquisition is deferred until the arena receives its first allocation that can benefit from pooling. This is important because arenas that never allocate memory do not cause a native pool to be created. Once acquired, the pool belongs exclusively to that arena. After that, suitable allocations are carved sequentially from the pool. Here is an illustration of carving:

64-byte pool

0          8         16    20                          63
+----------+----------+-----+-------------------- ... --+
| pointer  |   long   | int | remaining capacity  ...   |
+----------+----------+-----+-------------------- ... --+

Figure 1. Carving out small allocations.

If an allocation does not fit, or requires an alignment that the pool cannot guarantee, it follows the regular allocation path. This makes pooling an opportunistic optimization: large and unusually aligned allocations continue to work as before.

When the arena closes, the used portion of its pool is cleared and returned to the thread cache. If the cache is already full, the pool is released to the native allocator instead. Each arena that acquires a pool owns it exclusively, so nested confined arenas (i.e., where a try-with-resource block with a first arena contains yet another try-with-resource with a second arena) remain isolated. If no cached pool is available, the arena attempts to allocate a local pool of its own. On close, that pool is either added to the cache or released.

Virtual Threads

Creating a native-memory cache in every virtual thread would undermine their small footprint. Instead, a virtual thread temporarily acquires a pool from its current carrier thread.

The virtual thread is pinned only during the short cache handoff. The pool is then detached from the carrier cache and owned by the arena. Consequently, the virtual thread can migrate safely while the arena remains open. When the arena closes, the pool is returned to the cache of the (possibly different) carrier on which the virtual thread is then mounted.

This design avoids adding a native-memory cache to potentially millions of virtual threads while preserving exclusive pool ownership.

Preserving Safety

Pooling does not change the lifetime or accessibility rules of the Arena API.

After an arena closes, this is still true:

Before a pool becomes available for reuse, the portion used by the previous arena is zeroed out. Newly allocated pools are also zero-initialized. This preserves the API requirement that native segments returned by an arena contain zeroes and prevents data written by one arena from being exposed through a later arena.

The regular cleanup actions run before the pool is cleared and released. This ordering matters because a cleanup action may legitimately access the native region through a cleanup segment.

Available cached pools are deterministically released when their platform or carrier thread terminates. Pools belonging to open arenas are detached from the cache and are therefore not accidentally freed by carrier termination or virtual-thread migration.

Observable Behavior Changes

Although segment lifetime is unchanged, the lifetime of the underlying native allocation can be different. Closing a pooled arena can return its memory to the JDK cache instead of immediately invoking the native allocator. Applications and diagnostic tools must not rely on a one-to-one correspondence between arena allocation and native malloc or free calls.

Performance

Development benchmarks showed large reductions in latency for tiny confined allocations. The following table reports nominal speedups for allocations of 5 and 20 bytes:

Platform 5-byte allocation 20-byte allocation
Linux AArch64 12.3x 8.6x
Linux x64 6.8x 7.8x
macOS AArch64 18.6x 16.8x
Windows x64 16.8x 18.9x

Table 1. Nominal speedups across the tested platforms.

Each entry is the baseline JMH average-time latency divided by the pooled latency; higher is better. The benchmark configuration and raw results are available in OpenJDK PR #31365.

A separate benchmark run on an Apple M4 running macOS produced:

Benchmark                                   (size)  Mode  Cnt    Score   Error  Units
AllocTest.alloc_confined                         5  avgt   30    1.052 ± 0.027  ns/op
AllocTest.alloc_confined                        20  avgt   30    1.138 ± 0.023  ns/op
AllocTest.alloc_confined                       100  avgt   30   19.044 ± 0.322  ns/op
AllocTest.alloc_confined                       500  avgt   30   26.383 ± 0.714  ns/op
AllocTest.alloc_confined                      2000  avgt   30   31.292 ± 0.253  ns/op
AllocTest.alloc_confined                      8000  avgt   30   81.119 ± 0.957  ns/op
AllocTest.alloc_confined_no_pool                 5  avgt   30   15.862 ± 0.226  ns/op
AllocTest.alloc_confined_no_pool                20  avgt   30   18.139 ± 0.262  ns/op
AllocTest.alloc_confined_no_pool               100  avgt   30   18.856 ± 0.281  ns/op
AllocTest.alloc_confined_no_pool               500  avgt   30   26.168 ± 0.232  ns/op
AllocTest.alloc_confined_no_pool              2000  avgt   30   31.209 ± 0.314  ns/op
AllocTest.alloc_confined_no_pool              8000  avgt   30   80.840 ± 0.804  ns/op

Listing 1. Benchmarks on an M4 Mac.

For allocations above the default 64-byte pool size, performance remained broadly around parity because those allocations use the regular path. Virtual threads incur some additional carrier-handling overhead, but still showed substantial improvements over non-pooled allocation:

Benchmark                                   (size)  Mode  Cnt    Score   Error  Units
AllocTest.OfVirtual.alloc_confined               5  avgt   30    3.222 ± 0.087  ns/op
AllocTest.OfVirtual.alloc_confined              20  avgt   30    2.737 ± 0.021  ns/op
AllocTest.OfVirtual.alloc_confined             100  avgt   30   18.868 ± 0.199  ns/op
AllocTest.OfVirtual.alloc_confined             500  avgt   30   26.117 ± 0.240  ns/op
AllocTest.OfVirtual.alloc_confined            2000  avgt   30   31.407 ± 0.508  ns/op
AllocTest.OfVirtual.alloc_confined            8000  avgt   30   80.474 ± 0.676  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool       5  avgt   30   15.672 ± 0.090  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool      20  avgt   30   18.109 ± 0.250  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool     100  avgt   30   18.791 ± 0.163  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool     500  avgt   30   25.947 ± 0.174  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool    2000  avgt   30   31.206 ± 0.440  ns/op
AllocTest.OfVirtual.alloc_confined_no_pool    8000  avgt   30   79.711 ± 0.576  ns/op

Listing 2. Virtual thread benchmarks on an M4 Mac.

As always, microbenchmark speedups do not translate directly into equivalent whole-application speedups. The largest benefits are expected where arena creation, native allocation, and cleanup form a significant part of a short native operation.

Who Benefits?

This optimization is particularly relevant for:

Note that these benefits only apply to confined arenas and not the other arena types.

A Transparent Improvement

The most useful performance improvements are often those that do not require developers to choose a new API or rewrite working code. Pooled confined arenas preserve the deterministic lifetime, zero-initialization, and thread-confinement properties of Arena.ofConfined().

The difference is underneath: many small, short-lived allocations can now avoid repeated trips through the native allocator.

For FFM applications that use confined arenas as scoped scratch space, the existing code retains its safety properties while running considerably faster.

Further Reading

Here are some links if you are interested in learning more about pooled confined arenas: