WEKA NeuralMesh
by WEKA (WekaIO)
Microservices storage and memory fabric built for AI training and inference at exabyte scale
WEKA NeuralMesh is a storage and memory platform for AI workloads, built as distributed microservices with no central controller or metadata server. It serves POSIX, S3, NFS, SMB and NVIDIA GPUDirect Storage simultaneously over the same data, and extends GPU memory by holding KV cache on NVMe. It is aimed at enterprises and GPU cloud providers running large training and production inference fleets.
WEKA NeuralMesh is the storage and memory platform behind WEKA's AI infrastructure business, architected as distributed microservices rather than a conventional parallel file system, with no central controller and no metadata server, which is how it sustains microsecond latency at exabyte scale. A single dataset is addressable simultaneously through POSIX, S3, NFS, SMB and NVIDIA GPUDirect Storage with no duplication or gateway translation. Its most consequential AI feature is Augmented Memory Grid, which extends GPU memory by persisting KV cache to NeuralMesh-managed NVMe; benchmarked in production on Oracle Cloud Infrastructure H100 systems, WEKA reports 10x higher token throughput, 10x more concurrent users served and 7x more tokens per GPU against DRAM-based alternatives. NeuralMesh Axon runs converged directly on GPU servers, and a Kubernetes Operator with CSI drivers handles lifecycle and provisioning. On 22 July 2026 WEKA announced NeuralMesh 6, which it calls the largest software release in company history and which reaches general availability in the second half of 2026 at no additional cost to existing customers. It adds two-tier multi-tenancy - Composable Clusters for hardware isolation plus VPC-style virtual tenancy with per-tenant QoS, independent encryption keys and LDAP or Active Directory authentication - supporting over 1,000 isolated tenants per cluster and up to 50,000 logical tenants across composite deployments. It also brings a native S3 implementation handling 2,000-5,000 concurrent connections per node, metadata-first replication with on-demand hydration, and always-on data reduction carrying contractual guarantees up to 6x. Alongside it, WEKApod 3 appliances deliver 1.1 exabytes of effective capacity, 10.2 TB/s and 210 million IOPS per rack. WEKA serves 300+ customers including 12 of the Fortune 50.
The infrastructure or platform leader running a large GPU fleet - an AI-native company, a neocloud, or an enterprise AI factory - where GPUs sit idle waiting on data and the cost per token at sustained load is now the metric that matters.
Higher GPU utilization from one unified stack for training and inference, including KV cache offload to NVMe that WEKA benchmarks at 10x token throughput and 7x more tokens per GPU versus DRAM-based approaches on OCI H100 infrastructure.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Subscription, Contact for pricing
- Target Market
- CTOs, CIOs, Infrastructure Architects, ML Platform Engineers, Data Scientists, HPC Engineers
- Deployment
- Hybrid, Self-hosted, Multi-cloud, Cloud-first
- Founded
- 2013
- Headquarters
- Campbell, California, United States
- Team Size
- 201-500
- Customers
- 300+ customers including 12 of the Fortune 50; named users include Stability AI, Cohere, Midjourney, ElevenLabs, the Center for AI Safety, IREN, Applied Digital, NexGen Cloud and Yotta
Key Features
- ✓Microservices architecture with no central controller
Distributes metadata and data services across nodes to sustain microsecond latency at exabyte scale without a single-point metadata bottleneck.
- ✓Augmented Memory Grid
Persists KV cache to NVMe-backed NeuralMesh storage to extend GPU memory, benchmarked on OCI H100 at 10x token throughput versus DRAM alternatives.
- ✓Unified multi-protocol access
Serves POSIX, S3, NFS, SMB and NVIDIA GPUDirect Storage against the same physical blocks simultaneously, eliminating copies between protocol silos.
- ✓Two-tier multi-tenancy in NeuralMesh 6
Combines Composable Clusters for hardware isolation with virtual network tenancy, supporting 1,000+ isolated tenants per cluster and 50,000 logical tenants.
- ✓Always-on data reduction
Applies fingerprinting, similarity hashing, deduplication and compression by default with under 5% write overhead and contractual guarantees up to 6x.
- ✓NeuralMesh Axon converged deployment
Runs the storage fabric directly on GPU servers, collapsing a separate storage tier into the compute nodes already purchased.
- ✓Kubernetes Operator and NeuralMesh Observe
Automates cluster deployment and lifecycle via native CSI drivers, with SaaS multi-cluster dashboards, client diagnostics and Slack or PagerDuty alerting.
Capabilities
Use Cases
- •Large model training at scale
Foundation model teams keep thousands of GPUs saturated by removing the storage bandwidth ceiling that otherwise leaves expensive accelerators idle.
- •Production inference with KV cache offload
Augmented Memory Grid holds persistent KV cache on NVMe, raising concurrent user counts and tokens per GPU without adding accelerator hardware.
- •Multi-tenant GPU cloud services
Neocloud providers isolate customers with dedicated hardware or virtual tenancy, per-tenant QoS and independent encryption keys on shared infrastructure.
- •Cross-site dataset mobility
Metadata-first replication with on-demand hydration lets teams place workloads where GPUs are available rather than where the data already lives.
- •Consolidating AI data silos
One platform serves file and object workloads together, removing the duplicate copies that separate training, preprocessing and serving tiers usually require.
Ideal For
Best For
- ✓Feeding large GPU training clusters where storage bandwidth, not compute, is the bottleneck on utilization
- ✓Production LLM inference needing persistent KV cache offload to extend effective GPU memory and raise concurrency
- ✓GPU cloud and neocloud providers that must isolate many customers on shared hardware with per-tenant QoS and encryption
- ✓Mixed pipelines that need POSIX, S3, NFS, SMB and GPUDirect access to the same dataset without copying it between silos
- ✓Kubernetes-native AI platforms requiring CSI provisioning and operator-managed lifecycle at rack scale
Not Ideal For
- ✗Cost-sensitive buyers or general-purpose enterprise file storage - reviewers on G2 and Gartner Peer Insights consistently flag WEKA as expensive relative to alternatives
- ✗Cold archive, backup or bulk capacity tiers, since the NVMe-based architecture and high-speed networking requirement are wasted on infrequently accessed data
- ✗Teams that need NeuralMesh 6's multi-tenancy today, as that release is not generally available until the second half of 2026 and WEKApod 3 ships from autumn 2026
- ✗Small AI teams without a substantial GPU fleet, where the platform's economics and operational complexity cannot be justified
Integrations
Deployment
Market & Ratings
300+ customers including 12 of the Fortune 50; named users include Stability AI, Cohere, Midjourney, ElevenLabs, the Center for AI Safety, IREN, Applied Digital, NexGen Cloud and Yotta
Market Analysis
Pros
- ✓Exceptional peer validation - 4.9 out of 5 across 119 Gartner Peer Insights reviews, with 98% of reviewers saying they would recommend it, and repeated Visionary placement in Gartner's Magic Quadrant
- ✓Augmented Memory Grid addresses inference economics directly, with production OCI H100 benchmarks of 10x token throughput and 7x more tokens per GPU
- ✓One software stack spans training and inference across on-premises, cloud, converged-on-GPU and appliance deployments, avoiding separate platforms per workload
- ✓Contractual data reduction guarantee and a free NeuralMesh 6 upgrade for existing customers are concrete commercial commitments rather than marketing claims
Cons
- ✗Consistently described by reviewers on G2 and Gartner Peer Insights as expensive compared with similar storage products, which is the single most recurring complaint
- ✗The headline NeuralMesh 6 capabilities - native multi-tenancy, unified file-and-object stack, always-on reduction - are not generally available until the second half of 2026, and WEKApod 3 appliances only ship from autumn 2026, so much of the announcement is forward-looking
- ✗Performance claims including 10x token throughput, 6x data reduction and 267% capacity density are vendor benchmarks; the supporting commentary at launch came from WEKA's own partners rather than independent testing
- ✗Requires NVMe media and high-speed networking, so it is a poor fit for capacity or archive tiers and raises the entry cost for smaller deployments
- ✗Choosing the WEKApod appliance route ties capacity expansion to WEKA-designed hardware, reducing the commodity-server flexibility the software otherwise allows
Pricing
NeuralMesh Subscription
Contact for pricing
- ✓Software subscription scaled by capacity and performance tier
- ✓Deploy on-premises, in cloud, converged on GPU servers or via WEKApod appliance
- ✓Contractual data reduction ratio guarantee
- ✓NeuralMesh 6 upgrade included at no extra cost for existing customers
WEKA publishes no list pricing. NeuralMesh is sold as a subscription scaled by data capacity, performance requirements and support level, with enterprise support and integration typically quoted separately; configurable SKUs span from under 1 PB to over 100 PB in a single deployment. WEKA backs its data reduction ratio with a contractual guarantee that covers the software licence for free if targets are missed, which is unusually concrete. Reviewers on G2 and Gartner Peer Insights repeatedly describe the platform as expensive versus comparable storage, so total cost should be modelled against recovered GPU utilization rather than against price per terabyte.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 5 on domains other than weka.io.
- 1.hpcwire.com — weka launches neuralmesh 6 and wekapod as ai inference pushe
- 2.storagereview.com — wekas wekapod 3 breaks the single rack exabyte barrier as ne
- 3.unite.ai — weka launches neuralmesh 6 and wekapod 3 as ai infrastructur
- 4.gartner.com — neuralmesh
- 5.venturebeat.com — weka scores 140m to supercharge ai workloads with dynamic da
- 6.weka.io — neuralmeshvendor
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Tsuga
Bring-your-own-cloud observability that keeps telemetry, and its cost, inside your own AWS account
OpenObserve
Open-source, Rust-based observability on object storage — logs, metrics, traces and LLM telemetry in one binary
Coralogix
AI-native observability that queries logs, metrics and traces in place — index-free, in your own S3 bucket
SkyPilot
Run and scale AI workloads across Kubernetes, Slurm and 25+ clouds from one interface