
CoreWeave and Hudson River Trading Turn Rubin Into Wall Street Infrastructure
Hudson River Trading's Rubin deployment on CoreWeave shows quantitative finance is becoming an early market for tightly integrated AI supercomputing.
CoreWeave announced on August 20 that Hudson River Trading will build a next-generation research platform on CoreWeave Cloud, utilizing a combination of NVIDIA Vera Rubin NVL72 and HGX B200 systems. This deployment marks a pivotal moment in the latest AI news, signaling that top-tier quantitative finance firms are moving beyond experimental clusters and committing massive capital to tightly integrated, rack-scale AI supercomputing for their most critical research and production pipelines.
The agreement matters because it provides the clearest evidence yet that the financial sector is becoming an anchor market for NVIDIA’s most advanced and power-dense architectures. Hudson River Trading is a premier quantitative trading firm, meaning its daily operations rely on a punishing workload mix that includes sprawling research pipelines, constant model experimentation, and highly time-sensitive production decisions. By securing this workload, CoreWeave is proving that specialized AI cloud providers can capture the most demanding enterprise customers on Wall Street, outmaneuvering traditional hyperscalers by offering faster access to the bleeding edge of silicon.
The Intersection of Quantitative Finance and Supercomputing
Quantitative trading has always been an arms race defined by compute, but the nature of that compute is fundamentally changing. For decades, firms like Hudson River Trading relied on vast clusters of traditional CPUs to backtest strategies against historical market data, alongside field-programmable gate arrays to execute trades with microsecond latency. However, the explosion of alternative data and the increasing viability of large language models in financial forecasting have strained traditional infrastructure. Today, generating alpha requires processing unstructured data—ranging from global supply chain logs and satellite imagery to real-time central bank communications—at a scale that only specialized AI infrastructure can handle.
The deployment on CoreWeave Cloud illustrates a structural shift in how quantitative research is conducted. A modern financial research platform is no longer just a repository of historical tick data; it can also be an iterative engine for AI training and inference. HRT and CoreWeave have not disclosed dataset sizes, latency targets, model architectures, capacity, or commercial terms. What the announcement does establish is HRT's plan to use both Rubin NVL72 and HGX B200 systems for a next-generation research platform.
By leveraging CoreWeave Cloud, Hudson River Trading gains access to an environment optimized specifically for accelerated workloads. CoreWeave emphasizes infrastructure designed around the thermal, networking, and power requirements of NVIDIA’s densest racks. That specialization can matter for quantitative research, although the announcement provides no benchmark comparing this deployment with a general-purpose cloud or HRT's existing systems.
The inclusion of both HGX B200 and Vera Rubin NVL72 systems in this deployment highlights a sophisticated, multi-tiered approach to algorithmic research. The HGX B200 systems, based on the Blackwell architecture, are likely to serve as the workhorses for model experimentation and foundational AI training. These systems excel at the brute-force matrix multiplication required to train large language models and complex neural networks from scratch. Meanwhile, the Vera Rubin NVL72 represents the frontier of high-throughput inference and specialized processing, designed to handle the massive parameter counts and complex routing of next-generation AI models at unprecedented efficiency.
Dissecting the Vera Rubin Architecture
To understand the significance of this deployment, one must look closely at what NVIDIA has built with the Vera Rubin generation. NVIDIA explicitly describes Vera Rubin as a six-chip platform spanning CPU, GPU, networking, DPU, and interconnect components, rather than a standalone accelerator. This distinction is crucial for the Artificial Intelligence News cycle, as it redefines what buyers are actually purchasing when they invest in modern AI infrastructure.
The Vera Rubin platform integrates the Vera CPU with the Rubin GPU, binding them together with next-generation NVLink interconnects, advanced ConnectX networking interfaces, and BlueField Data Processing Units. This holistic design acknowledges that the bottleneck in modern AI is rarely the raw compute capability of the GPU itself; rather, it is the speed at which data can be moved into, out of, and between the processing cores. By engineering all six components to operate as a single, cohesive supercomputer, NVIDIA has drastically reduced the latency and power overhead associated with data movement.
| Component Category | NVIDIA Rubin Platform Element | Function in Quantitative Finance Workloads |
|---|---|---|
| Central Processing | Vera CPU | Handles complex data preprocessing, orchestrates market data ingestion, and manages sequential logic in trading algorithms before handing off parallel tasks. |
| Graphics Processing | Rubin GPU | Executes massive matrix multiplications for AI training and high-throughput inference for real-time predictive modeling. |
| Interconnect | NVLink Switch | Enables unified memory access across the entire NVL72 rack, allowing massive models to operate as if on a single giant GPU without PCIe bottlenecks. |
| Network Interface | ConnectX | Facilitates ultra-high-bandwidth, low-latency communication between different racks in the CoreWeave data center for distributed training jobs. |
| Data Processing | BlueField DPU | Offloads infrastructure tasks like security, storage routing, and network virtualization, ensuring the GPUs are dedicated entirely to alpha generation. |
For a firm like Hudson River Trading, this architectural integration is a massive advantage. In quantitative research, models are often so large that they cannot fit into the memory of a single GPU, or even a single server node. They must be distributed across dozens or hundreds of chips. The NVL72 configuration acts as a massive, single-rack supercomputer where 72 GPUs share a unified memory space via the NVLink switch fabric. This allows researchers to train and run inference on colossal models without the severe latency penalties typically incurred when data has to traverse traditional Ethernet or InfiniBand networks between separate servers.
Networking and the Data Processing Unit
The inclusion of the BlueField DPU in the Vera Rubin platform is particularly relevant for financial workloads. In a traditional setup, the CPU spends a significant portion of its cycles managing network traffic, handling storage requests, and enforcing security protocols. In a high-frequency trading or quantitative research environment, these overheads introduce unpredictable latency—the enemy of algorithmic efficiency.
The DPU offloads these infrastructure tasks from the Vera CPU, creating a secure, isolated, and highly deterministic environment for the primary workload. For Hudson River Trading, this means that market data feeds can be ingested, decrypted, and routed directly to GPU memory with minimal CPU intervention, a technology NVIDIA calls GPUDirect. This streamlined data path is essential when processing the firehose of alternative data required to feed modern AI agents and predictive models in real time.
CoreWeave's Strategic Position in Specialized Cloud
CoreWeave’s ability to secure this contract underscores its rising dominance in the specialized AI cloud market. The company’s second-quarter materials list Hudson River Trading among its customers and heavily emphasize the surging demand for specialized AI cloud capacity. While traditional hyperscalers like Amazon Web Services, Google Cloud, and Microsoft Azure offer immense scale, their infrastructure is inherently generalized to support everything from simple web hosting to enterprise databases.
CoreWeave, by contrast, has built its entire business model around the unique physical and networking demands of NVIDIA’s highest-end hardware. The Vera Rubin NVL72 racks are incredibly dense and run exceptionally hot, requiring advanced liquid cooling solutions and power delivery mechanisms that many legacy data centers simply cannot support without massive retrofitting. CoreWeave has purpose-built its newer facilities to accommodate these exact specifications, allowing it to deploy and scale these systems faster than competitors burdened by legacy infrastructure.
This strategic positioning is highly attractive to quantitative finance firms. In the race to develop the most accurate predictive models, time-to-compute is a critical metric. Waiting six months for a hyperscaler to provision the latest hardware can mean missing an entire market cycle. CoreWeave has stated that Rubin-based systems will enter its cloud in the second half of 2026, offering Hudson River Trading a rapid path to deployment for its next-generation research platform.
Analyzing the DeepSeek R1 Efficiency Benchmark
One of the most striking aspects of this deployment is the performance context provided by CoreWeave. The cloud provider has published a benchmark claim stating that the Vera Rubin NVL72 delivers up to ten times more tokens per second per megawatt than the Blackwell GB200 NVL72 on a specified DeepSeek R1 workload. This is a staggering metric that requires careful unpacking, as it highlights a fundamental shift in how AI infrastructure is evaluated.
First, it is vital to recognize that this is a vendor benchmark, not an independent, universal result. Performance in AI workloads is highly dependent on the specific model architecture, batch size, sequence length, and optimization techniques used. However, the choice of the DeepSeek R1 workload for this benchmark is highly instructive. DeepSeek R1 is known as a complex, reasoning-heavy model that likely utilizes a Mixture of Experts architecture.
Mixture of Experts models are notoriously difficult to run efficiently because they require routing individual tokens to specific sub-networks (experts) distributed across multiple GPUs. This routing demands immense memory bandwidth and ultra-fast interconnects. The fact that the Rubin NVL72 excels so dramatically on this specific workload suggests that the NVLink switch fabric and the unified memory architecture of the 72-GPU rack are functioning exactly as NVIDIA intended, eliminating the communication bottlenecks that throttle Mixture of Experts models on older architectures.
Furthermore, the metric itself—tokens per second per megawatt—reveals the new reality of data center economics. We are no longer strictly measuring raw floating-point operations per second. Power has become the ultimate limiting factor in AI scaling. Data centers are hitting the limits of local utility grids, and the cost of electricity is becoming a dominant factor in the total cost of ownership for AI infrastructure.
By framing the benchmark around power efficiency, CoreWeave is directly addressing the primary concern of enterprise buyers. For a firm like Hudson River Trading, a 10x improvement in power efficiency means they can run vastly more complex models, or run the same models ten times faster, without exceeding their allocated power budget in the data center. This efficiency translates directly into a competitive advantage, allowing for deeper historical backtesting and more granular real-time inference within the same operational footprint.
flowchart LR
A[Market and alternative data] --> B[HRT research pipelines]
B --> C[CoreWeave AI cloud]
C --> D[HGX B200 systems]
C --> E[Vera Rubin NVL72]
D --> F[Training and simulation]
E --> G[High throughput inference]
F --> H[Quantitative research decisions]
G --> H
The Role of HGX B200 in the Research Pipeline
While the Vera Rubin NVL72 commands the headlines with its massive rack-scale integration and inference efficiency, the inclusion of HGX B200 systems in the Hudson River Trading deployment is equally significant. The HGX B200, based on the Blackwell architecture, remains an incredibly powerful platform for AI training.
In a quantitative research environment, the workflow is rarely linear. Quants constantly hypothesize new market signals, requiring them to train new models or fine-tune existing ones on fresh datasets. This experimentation phase is highly iterative and requires massive parallel compute to churn through petabytes of historical data. The HGX B200 systems are perfectly suited for this task, offering the raw computational horsepower needed to compress training times from weeks to days, or days to hours.
Once a model has been trained and validated on the HGX B200 clusters, it can be quantized, optimized, and deployed to the Vera Rubin NVL72 racks for high-throughput inference and real-time simulation. This bifurcated approach allows Hudson River Trading to optimize its infrastructure spend, using the right architectural tool for the specific phase of the research pipeline. It prevents the highly specialized, interconnect-heavy Rubin racks from being bogged down by brute-force training jobs that do not fully utilize the NVLink switch fabric, ensuring maximum return on investment across the entire CoreWeave deployment.
Workload Realities for Wall Street AI Agents
The integration of these advanced systems into Wall Street infrastructure points to a broader evolution in how financial firms utilize AI tools. We are moving rapidly past the era where large language models were used merely for text summarization—parsing earnings call transcripts or summarizing analyst reports. The new frontier involves deploying autonomous AI agents capable of complex reasoning, multi-step planning, and direct interaction with algorithmic trading systems.
In the context of quantitative finance, an AI agent might be tasked with monitoring a specific sector of the global economy. This agent would continuously ingest real-time news feeds, satellite data regarding shipping lanes, and social media sentiment. Using a reasoning-heavy model similar to the DeepSeek R1 used in CoreWeave's benchmark, the agent would synthesize this unstructured data, identify emerging patterns, and generate a probabilistic forecast of asset price movements.
Crucially, these generative AI systems must operate with strict latency bounds. A forecast generated five minutes after a market-moving event is entirely useless to a high-frequency trading algorithm. The Vera Rubin NVL72, with its massive memory bandwidth and integrated networking, provides the substrate necessary to execute these complex, agentic workflows in near real-time. By processing tokens at unprecedented speeds, the infrastructure ensures that the AI agents can complete their reasoning loops and deliver actionable signals to the execution algorithms before the market opportunity evaporates.
This represents a paradigm shift in alpha generation. Traditional quantitative models rely heavily on structured, numerical data—price, volume, order book depth. Generative AI and LLMs unlock the vast universe of unstructured data, allowing quants to model the actual drivers of market sentiment and economic activity. However, extracting signal from this noise requires immense computational power, which is exactly why firms like Hudson River Trading are securing dedicated capacity on platforms like the NVIDIA Vera Rubin.
Operational Tradeoffs in High-Density Deployments
Despite the massive performance gains promised by the Rubin architecture, deploying infrastructure at this scale involves significant operational tradeoffs. The transition from traditional CPU clusters to rack-scale AI supercomputers fundamentally alters the economics and risk profile of a quantitative research platform.
| Workload Type | Primary Hardware | Performance Metric | Deployment Phase |
|---|---|---|---|
| Model Experimentation | HGX B200 | Time-to-train, raw FLOPS | Research and Development |
| Historical Backtesting | HGX B200 | Batch processing throughput | Strategy Validation |
| Agentic Reasoning | Vera Rubin NVL72 | Tokens per second per megawatt | Pre-production Simulation |
| Real-time Inference | Vera Rubin NVL72 | End-to-end latency, memory bandwidth | Production Deployment |
The most immediate tradeoff is cost. Securing dedicated capacity on the latest NVIDIA hardware via CoreWeave requires a massive capital commitment. While CoreWeave abstracts away the physical data center management, the cloud consumption costs for NVL72 racks will be astronomical compared to legacy compute. Financial firms must carefully calculate whether the incremental alpha generated by these advanced models justifies the exponential increase in infrastructure spend.
Furthermore, there is the issue of vendor lock-in. By building their next-generation research platform around the specific capabilities of the NVIDIA 6-chip platform and the CoreWeave cloud environment, Hudson River Trading is deeply integrating its proprietary research pipelines with a specific hardware and software stack. NVIDIA's CUDA ecosystem and its proprietary NVLink interconnects create a highly optimized but closed environment. If a competing silicon provider were to release a more efficient architecture, migrating the research platform would require a massive engineering effort to rewrite custom kernels and re-optimize data pipelines.
Power and cooling also remain persistent challenges, even in a cloud environment. While CoreWeave manages the physical infrastructure, the sheer density of the Rubin NVL72 racks means that compute capacity is strictly bounded by the facility's ability to deliver power and reject heat. If a quantitative firm needs to suddenly scale its inference capacity during a period of extreme market volatility, they may find that even their dedicated cloud provider cannot instantly spin up additional NVL72 racks due to hard physical constraints in the data center.
Evidence Limits and Vendor Claims
As with any major infrastructure announcement in the Artificial Intelligence News space, it is crucial to clearly bound the evidence and separate verified facts from marketing projections. CoreWeave has confirmed that Hudson River Trading is building this platform, and NVIDIA has detailed the architectural specifications of the Vera Rubin platform. However, the operational realities of this deployment remain to be proven in the wild.
The benchmark claim of up to ten times more tokens per second per megawatt for the Vera Rubin NVL72 compared to the Blackwell GB200 NVL72 must be treated with caution. As noted, this is a vendor-supplied metric based on a specific, highly demanding workload (DeepSeek R1). While it demonstrates the theoretical peak efficiency of the architecture on Mixture of Experts models, it does not guarantee a 10x performance improvement across all of Hudson River Trading's diverse research pipelines. Many traditional quantitative models and smaller neural networks may not fully saturate the NVLink fabric, resulting in much more modest efficiency gains.
Additionally, the timeline for this deployment is subject to the realities of hardware manufacturing and supply chain logistics. CoreWeave has stated that Rubin-based systems will enter its cloud in the second half of 2026. However, integrating these complex, liquid-cooled racks into existing data center topologies is a non-trivial engineering challenge. Delays in silicon yields, networking components, or cooling infrastructure could push the actual availability of production-ready capacity further into the future. Buyers and builders must monitor CoreWeave's actual deployment velocity in Q3 and Q4 of 2026 to verify that these systems are scaling as promised.
The Future of Algorithmic Market Infrastructure
The partnership between CoreWeave and Hudson River Trading serves as a leading indicator for the broader financial services industry. The deployment of the Vera Rubin NVL72 is not just an infrastructure upgrade; it is a fundamental re-architecture of how market intelligence is generated. We are witnessing the convergence of high-performance computing, generative AI, and quantitative finance into a single, tightly coupled discipline.
Looking ahead, builders of AI tools and enterprise buyers should expect the definition of a "research platform" to shift dramatically. The era of assembling disparate clusters of CPUs and standalone GPUs over standard Ethernet is ending for top-tier firms. The future belongs to integrated, rack-scale supercomputers where the CPU, GPU, DPU, and network fabric are co-designed to eliminate data movement bottlenecks.