Today Azure Storage introduces in preview a new List Blobs optimization that accelerates listing operations by up to 25x with up to 15x lower client-side CPU utilization allowing customers to return millions of objects per second in List Blobs results.
The list results are now returned in Apache Arrow format, a highly optimized and compact columnar response format that allows more efficient parsing and reduced client-side CPU utilization. Clients can now parallelize object list operations efficiently across multiple concurrent requests, while maintaining the same strong consistency of listing results that applications require. This increased performance and lower client CPU utilization is delivered within the existing List Blobs API via just a simple header change and is automatically invoked when using the updated Azure Blob Storage SDK’s.
Why we built this
Object storage was originally designed to provide low-cost, resilient data access through simple REST APIs. Early systems contained only a few million objects and were listed infrequently, so listing performance was not a priority.
Cloud, big data, mobile, and cloud-native computing steadily increased data volumes and access demands. Since 2020, foundation models, LLMs, and large GPU fleets have pushed object storage to trillions of objects, making it an active data layer for AI.
Modern AI and analytics workloads operates at the scale of trillions of objects. These workloads must repeatedly discover and inventory vast datasets for pre-training, analytics, fine-tuning, and inference, making frequent listing operations a significant source of storage-system pressure and client-side CPU consumption.
Downstream AI training and data analytics jobs are gated by the time required to enumerate immense datasets, while parsing the results consumes client-side CPU that could otherwise run the workload.
To meet these rapidly growing demands of AI workloads, we’ve built the next generation of Azure Blob Storage listing capabilities that deliver the performance required at this massive scale. Next, we will dive deeper into the details of how List Blobs performance and scalability have been accelerated and how simple it is for customers to take advantage of this new level of performance.
Faster, more efficient, listing results with lower client CPU utilization
The Apache Arrow format was selected as the most efficient new way to deliver the performance and scalability gains required for efficiently listing millions of objects per second. Apache Arrow is an open-source, compact columnar format that returns List Blobs results in an optimized response roughly one third the size of the XML format used by most cloud object-storage listing APIs, making parsing faster and easier.
The new format is enabled with a simple request-header change. The List Blobs API then packages and accelerates results automatically while preserving strong consistency, so newly written objects remain immediately visible. Because Apache Arrow is compact and efficient to parse, client-side CPU utilization per object listed decreased by up to 15x, freeing compute resources for the workload itself.
Submitting multiple List Blobs requests concurrently further improves performance, delivering up to a 25x increase in enumeration speed with a single request-header change, while preserving schema consistency and compatibility with existing XML responses.
This new List Blobs performance enhancement via Apache Arrow remains an additive, opt-in extension of the existing List Blobs API, not a replacement. The existing XML-based List Blobs API remains unchanged, and current clients not adopting the new header continue to work with no breaking changes, but without the performance and lower CPU-utilization benefits.
The next figure shows the comparison of a parallelized listing operation using Apache Arrow format compared with the XML baseline on 16 clients and 48 threads per client.
rclone accelerates listing of 100K objects from 24 seconds to 1.1 seconds
rclone is a popular command-line program to manage, copy, move and replicate files on and between cloud storage destinations. rclone is a widely used open-source tool for moving and syncing data across almost all types of cloud storage.
Listing is a critical part of rclone’s synchronization workflow. To perform these data management operations at scale for large numbers of objects, rclone must enumerate the objects that are intended to be copied, moved, or replicated. For example, before and after syncing Azure Blob Storage containers, rclone performs a large-scale listing to compare the source and target states.
After updating rclone to incorporate the simple header change for the enhanced Apache Arrow-based List Blobs API calls, rclone was able to reduce their List Blobs time-to-completion by 21.7x from 24 seconds to 1.1 seconds on 100K object datasets.
The chart below shows rclone’s measured wall clock time for listing a single container with 100,000 entries (columns, left axis) alongside the resulting speedup versus the classic XML path (line, right axis).
Two things stand out in the results:
- Apache Arrow is an immediate win on its own: With no parallelism at all, sequential Arrow listing is 3.5× faster than the current XML-based List Blobs results, reducing listing operation time to completion from 23.9 seconds to 6.9 seconds .
- Parallel enumeration compounds the gains: Throughput climbs steadily with concurrency, reaching a 21.7× speedup at a parallelism of 30 and completing the same listing of 100,000 entries in just 1.1 seconds.
The accelerated List Blobs results are returned with the same consistency using the same List Blobs API but now returned much faster with no breaking changes for existing clients. That is exactly the outcome we set out to deliver with this new capability.
In rclone’s own words about these results:
” rclone has to list containers before it can sync them; with Apache Arrow and parallelism enabled this will make a sync of a directory with millions of files get going 20x faster. The Azure Storage team has been very responsive to our feedback during the preview which made the integration straightforward. The new Go SDK works very well and required very few code changes. Our Azure Blob Storage users are going to love this!”
Nick Craig-Wood, rclone Lead Developer
To get started, you can download rclone from the official rclone website.
If you are running rclone v1.74.0 or later you can enable the Apache Arrow listing with the –azureblob-use-arrow-list flag and enable listing parallelism with –azureblob-list-parallelism. As described in the testing, if you set “–azureblob-list-parallelism 30” this will get you the most performance listing from Azure with Arrow listing also enabled.
A sincere thank you to the rclone community for adopting List Blobs with Apache Arrow early and sharing such clear, quantified results. Feedback like that is invaluable as we advance toward general availability. Happy listing rcloners!
How to get started
At the REST layer, the Arrow response is negotiated with the Accept: application/vnd.apache.arrow.stream header on a minimal x-ms-version of 2026-06-06 or later. This will return a response content that will be an Apache Arrow IPC stream that can be decoded and used to instantiate a RecordBatchStreamReader using Apache Arrow SDKs in any language.
An example of decoding from Rest API response is provided below using Apache Arrow Python SDK for decoding:
table = pa.ipc.open_stream(resp.content).read_all()
print(table.schema)
print(“\nrows:”, table.num_rows, ” columns:”, table.num_columns)
Name: string not null
Creation-Time: timestamp[s]
Last-Modified: timestamp[s]
BlobType: string
ResourceType: string not null
Etag: string
Content-Length: uint64
Content-Type: string
Content-MD5: string
AccessTier: string
AccessTierInferred: bool
LeaseState: string
LeaseStatus: string
ServerEncrypted: bool
— schema metadata —
NumberOfRecords: ‘100’
NextMarker: ”
rows: 100 columns: 14
import pandas as pd
df = table.to_pandas()
df[[“Name”, “BlobType”, “Content-Length”, “size_mb”, “AccessTier”]].head(6)
|
Name |
BlobType |
Content-Length |
AccessTier |
|
|
0 |
train_chunk10_shard1.jsonl.zst |
BlockBlob |
215 |
Hot |
|
1 |
train_chunk10_shard10.jsonl.zst |
BlockBlob |
216 |
Hot |
|
2 |
train_chunk10_shard2.jsonl.zst |
BlockBlob |
215 |
Hot |
|
3 |
train_chunk10_shard3.jsonl.zst |
BlockBlob |
215 |
Hot |
|
4 |
train_chunk10_shard4.jsonl.zst |
BlockBlob |
215 |
Hot |
|
5 |
train_chunk10_shard5.jsonl.zst |
BlockBlob |
217 |
Hot |
This new accelerated performance for List Blobs listing via Apache Arrow can be transparently enabled on Python, Java, .NET, C++ and Go SDKs with a simple option in the container listing function. Enabling via Azure Blob SDKs allows the performance benefits to be achieved without changing the listing interface and returned object formats in the Azure Storage SDK.
This new accelerated performance for List Blobs is available in public preview across the Azure Storage client libraries listed below. To evaluate the capability, use the corresponding minimum preview version for your preferred language.
|
SDK |
Minimum version (preview) |
|
.NET |
12.30.0-beta.1 |
|
Python |
12.31.0b1 |
|
Java |
12.36.0-beta.1 |
|
Go |
v1.8.1-beta.1 |
|
C++ |
12.19.0-beta.1 |
|
JavaScript |
12.34.0-beta.1 |
To get started, update to a preview enabled SDK for your language and start testing the new feature with our samples, or add the new listing options to your existing REST calls
Parallelizing listing operations unlocks the double-digit performance gains described in this article. The optimal strategy depends on the namespace layout and distributes listing requests across multiple threads. The two most common approaches are:
- Use delimiter parameter to recursively fan out additional threads for each BlobPrefix, walking the namespace.
- Partition the namespace with startFrom and endBefore, then process the ranges across multiple threads. This is the approach used by rclone.
The best approach depends on the specific namespace layout and can be optimized by tuning both the algorithm and the level of parallelism.
Limitations
This new Apache Arrow-powered performance optimization for List Blobs is today supported on flat namespace (FNS) Azure Blob Storage accounts. In scenarios where the account is Hierarchical Namespace (HNS) enabled, if the List Blobs REST API is called with the new header, it will return a 409 (Conflict) error code. This error code can be used to fallback on the client application to standard XML listing.
Public preview is where your input shapes the product. Try it on your largest containers, tell us what you measure, and let us know what would make it even better through this form.
References


