[RFC v4 2/3] lib: add fastmem library
Mattias Rönnblom
hofors at lysator.liu.se
Thu Aug 27 07:23:40 CEST 2026
On 8/13/26 13:32, Morten Brørup wrote:
> Hi Mattias,
>
> It would be great to get fastmem into DPDK 26.11.
>
At this point, it doesn't look like I have any time to spend on this
contribution in the near future.
Also, I would like to get more actual application use of this library
before its inclusion.
I started to investigate if you can fit a slab allocator as one of the
allocation paths in rte_malloc.h, and I think that might be doable as
well. There's a fair bit of subtle issues there to solve first though,
and that's a more ambitious project. My next RFC would probably have
been an attempt at this.
In case fastmem is accepted as a separate, standalone library it doesn't
need to be shipped in one of the ABI-breaking releases. So 27.03 should
be fine.
Something that would help is a tech board or mem subsystem maintainer
indication of if a separate API/library is the way to go as a first
step, or if we should do what Stephen Hemminger has argued for and not
introduce a new memory allocator API but rather just fix the heap right
away.
The latter is higher effort, higher risk, but also higher reward.
> I have reviewed the patch for features and usage. Not for correctness.
>
Thanks. I will review this later, when I pick up fastmem work again.
Sorry for the late reply.
> Here are my comments, in somewhat random order.
>
> -
>
> Export symbols for 26.11, not 24.11.
>
> -
>
> Generally, consider setting rte_errno and returning -1, instead of returning -ESOMETHING.
> Then, e.g. in bin_alloc_one(), if slab_acquire() returns NULL, grow_socket() (called by slab_acquire()) has already set rte_errno.
>
> -
>
> +/*
> + * Slab header at offset 0 of each 2 MiB slab. Either free (linked
> + * via next_free) or assigned to a bin (linked via list).
> + */
> +struct fastmem_slab {
> + struct fastmem_bin *bin;
> + void *free_head;
> + uint32_t free_count;
> + uint32_t n_slots;
> + struct fastmem_slab *next_free;
> + TAILQ_ENTRY(fastmem_slab) list;
> + rte_iova_t iova_base;
> +} __rte_aligned(FASTMEM_SLAB_HEADER_SIZE);
>
> I wonder if rte_fastmem_virt2iova() is going to be a hot function?
> Then it should be inline, and the iova_base field should be first in the fastmem_slab for easier access (and possibly faster on architectures without "load with offset" instructions).
>
> The rte_fastmem_virt2iova() function can be inline without exposing fastmem internals by using opaque pointer magic.
> The opaque pointer magic should be protected by a bunch of static_assert()'s in rte_fastmem.c, mentioning that rte_fastmem_virt2iova() depends on these.
> Something like:
>
> static inline rte_iova_t
> rte_fastmem_virt2iova(const void *ptr)
> {
> // struct fastmem_slab *slab;
> void *slab;
>
> // slab = slab_of((void *)(uintptr_t)ptr);
> // slab = (uintptr_t)ptr & ~(uintptr_t)FASTMEM_SLAB_MASK;
> slab = ((uintptr_t)ptr & ~(uintptr_t)((1 << 21) - 1));
> // rte_fastmem.c: static_assert(FASTMEM_SLAB_MASK == (1 << 21) - 1),
> "Mismatch in rte_fastmem_virt2iova()");
>
> // return slab->iova_base + ((uintptr_t)ptr - (uintptr_t)slab);
> return *(rte_iova_t *)slab + ((uintptr_t)ptr - (uintptr_t)slab);
> // rte_fastmem.c: static_assert(offsetof(struct fastmem_slab, iova_base) == 0,
> "Mismatch in rte_fastmem_virt2iova()");
> }
>
> -
>
> Generally, statistics counting should be build-time configurable enable/disable for performance.
> It's acceptable to keep the variables in the structs, and just not update them.
> Consider (with statistics disabled) removing the statistics variables from the fastmem_cache structure, so the objects at the bottom of the stack are in the same cache line as the count/capacity fields. (If it makes any practical difference.)
>
> -
>
> +struct fastmem_cache {
> + uint32_t count;
> + uint32_t capacity;
> + uint32_t target;
> + uint64_t alloc_cache_hits;
> + uint64_t alloc_cache_misses;
> + uint64_t alloc_nomem;
> + uint64_t free_cache_hits;
> + uint64_t free_cache_misses;
> + void *objs[];
> +} __rte_cache_aligned;
>
> "target" is a shadow of "capacity / 2".
> It is rarely used without also accessing "capacity".
> Consider dropping it, and just using capacity / 2.
> It might even improve performance in some locations, where "capacity" is already loaded into a CPU register or will be used shortly thereafter. (My performance comment is pure speculation, not measured.) (For reference, the mempool lib uses cache->size / 2.)
>
> -
>
> Initialization should be done at the proper DPDK startup stage, not lazily at all fastmem functions.
> In the fastmem functions, calling fastmem_assure() should be replaced by a simple RTE_ASSERT(fastmem != NULL).
>
> I was wondering how do other libs handle the situation where a secondary process depends on init performed by the primary function?
> It looks like they do something similiar to fastmem_assure(). This sucks!
> That's an important feature gap in DPDK, which I already registered as a bug on the wishlist: https://bugs.dpdk.org/show_bug.cgi?id=1913
>
> Until we get that wish fixed, we should make eal depend on fastmem, and call the fastmem init function from inside rte_eal_init().
>
> -
>
> Do you foresee any other flags than RTE_FASTMEM_F_ZERO?
> I'd prefer not passing a flags parameter to the fastmem public alloc APIs, and instead expose two variants of the alloc APIs.
> It's OK to pass around the zero flag (or a boolean) internally; it will be optimized away because it is known at build time from the two alloc API variants.
>
> -
>
> A fastmem handle "rte_fastmem_handle_t" is not an object handle, so the name could be misleading.
> It's rather a sort of accelerator or helper for faster alloc/free.
>
> Also, it doesn't support SOCKET_ID_ANY (-1).
> Would it be beneficial to support SOCKET_ID_ANY, so the socket_id part gets resolved when using the fastmem handle?
> Or would that defeat the performance gain of fastmem handles?
> And are there relevant use cases for fastmem handle supporting SOCKET_ID_ANY, or am I pointlessly feature creeping?
>
> -
>
> +#define FASTMEM_MIN_SIZE ((size_t)1 << FASTMEM_MIN_CLASS_LOG2)
> +#define FASTMEM_MAX_ALLOC_SIZE ((size_t)1 << FASTMEM_MAX_CLASS_LOG2)
>
> Min and Max names should be similar.
>
> -
>
> rte_fastmem_cache_flush():
> The function name could indicate that it operates on the current lcore only.
> Makes code using the function easier to read, for reviewers unfamiliar with the fastmem lib.
>
> -
>
> In grow_socket():
> + if (socket->n_memzones == FASTMEM_MAX_MEMZONES_PER_SOCKET) {
> + FASTMEM_LOG(ERR,
> + "reached per-socket memzone cap (%u) on socket %d",
>
> cap -> limit
>
> -
>
> In bin_push_locked():
>
> + struct fastmem_slab *slab = (struct fastmem_slab *)
> + ((uintptr_t)obj & ~(uintptr_t)FASTMEM_SLAB_MASK);
>
> -> slab_of()
>
> -Morten
More information about the dev
mailing list