
Singapore – On August 24, next-gen AI infrastructure platform B.AI hit a major milestone: cumulative token throughput across its sitewide free-access campaign officially crossed the 2 trillion threshold. The record-breaking run follows a series of aggressive, developer-first plays that have turned heads across the industry.
While rising prices from leading model providers have sparked cost anxiety across the ecosystem, B.AI took an opposite approach. The momentum kicked off on August 17, when B.AI opened up free access to DeepSeek V4 Flash—shattering platform throughput records at 220 billion tokens in a single day. That was followed by sitewide free access to Tencent Hy3, DeepSeek-V4-Flash-Vision-Exp, and Xiaomi MiMo-V2.5. The resulting traffic surge didn’t just showcase developer demand; it proved B.AI’s underlying architecture can comfortably handle massive, concurrent production loads at scale.
Anchored in its strategy as a global compute distribution hub, B.AI is stepping up as an industry disruptor to unleash the full efficiency of underlying infrastructure. It guarantees enterprise-grade availability while driving costs down by up to 90% through its hybrid API model featuring direct official routing for stability alongside budget-optimized custom pipelines. Backed by seamless Web2 and Web3 dual-rail payment integrations and ongoing user rebates, B.AI is building a seamless pipeline for truly accessible AI compute.
B.AI Unveils Super Compute Distribution Hub to Overhaul AI Infrastructure via Tiered APIs and Smart Routing
As compute demands become increasingly specialized, B.AI is cementing its status as the foundational infrastructure for next-gen AI development. To solve the exponential inference costs tied to AI Agent workloads, B.AI is moving beyond basic API aggregation—evolving into a high-throughput distribution hub that dynamically routes compute across global networks.

To back this vision, B.AI introduced a dual-tier API access model, Official and Custom Provider, offering developers flexible compute options that maximize cost efficiency and lower the cost barrier for building the foundation of tomorrow’s AGI.
Designed for core production environments and complex inference tasks, B.AI’s Official channel features direct API connections, guaranteeing maximum platform availability. Leveraging massive economies of scale, B.AI passes raw volume discounts directly to developers—offering baseline price cuts from 10% up to 40% off market rates. This allows enterprise teams to slash base compute overhead while securing iron-clad operational stability.

For non-critical workflows that prioritize cost control over absolute uptime, B.AI introduces the Custom Provider option. Developers can route traffic directly through vetted third-party vendors—including Mix, Nebula, and OL Station—at live-discounted rates. With seven discount tiers available, rates can plunge as low as 90% off standard pricing. This dual-option approach gives users total routing granularity, slashing total compute spend to the absolute floor.

Beyond rock-solid, developer-facing API infrastructure, B.AI streamlines everyday frontend user workflows with its native Auto mode embedded in the Chat interface. Tailored for natural language chat, text analytics, and daily business tasks, the system analyzes user intent per prompt and dynamically routes queries to the optimal underlying model in real time. This eliminates wasted capacity while making enterprise-grade AI models easier to use for non-technical users.
Crossing the 2-Trillion Token Milestone in 7 Days: How B.AI Is Democratizing Compute Through Radical Subsidies
B.AI has never wavered from its core thesis: building the definitive hub for accessible, high-performance compute. By consistently rolling out zero-cost campaigns for flagship foundational models and aggressively discounting API pipelines, B.AI delivers real, bottom-line savings to lower the barrier for enterprise AI deployment—earning widespread adoption and market validation along the way.
When DeepSeek recently announced price hikes, B.AI tapped its deep resource aggregation and ecosystem liquidity to make a bold market response. On August 17, B.AI made DeepSeek V4 Flash completely free for a limited time period, opening up both Web and API access so enterprises could run production-grade AI workloads with zero overhead.
This contrarian move sent shockwaves through the market. Within 24 hours of launch, B.AI’s core performance metrics shattered all-time highs across the board, with daily token volume surging past the 220-billion mark.
The momentum didn’t stop there. On August 21, Tencent’s highly anticipated Hy3 model went live on B.AI—fully free and open to all users. The very next day, on August 22, DeepSeek-V4-Flash-Vision-Exp followed suit, bringing cutting-edge visual reasoning into the zero-cost tier to further lower the compute cost for enterprise multimodal pipelines.
As access expanded, developer demand exploded. On August 24, B.AI hit a major milestone: cumulative token throughput across its sitewide free-access campaign officially crossed the 2 trillion threshold. Beyond setting new records for developer traction, these raw numbers deliver concrete proof of B.AI’s resilience, scalability, and dynamic load-balancing power under sustained, ultra-high concurrency, which is engineered into its core infrastructure.

In fact, B.AI has consistently operated at an aggressive, high-frequency pace of delivering value back to the community. Prior to this, the platform continually shared value with the market through major initiatives like limited-time free access to MiniMax M3 and Qwen-3.8 MAX, the exclusive 10% discount on GLM 5.3, and generous top-up bonuses. But this is just the beginning. Moving forward, B.AI will continue to scale its ecosystem perks, rolling out multi-tiered free access events and deep-discount campaigns. Working side by side with developers, B.AI remains committed to driving down AI deployment costs with a steady stream of subsidized compute.
This expanded commitment goes beyond standalone model promotions—it runs straight through core daily API operations. Recently, B.AI’s official API services underwent a major upgrade, expanding steep price discounts across mainstream foundation models. While guaranteeing direct-from-source stability, B.AI delivers significantly more cost-efficient compute configurations for enterprises and builders alike.
To serve a global, diverse developer ecosystem, B.AI has fully unified Web2 and Web3 payment rails, breaking down financial barriers worldwide. The platform integrates traditional fiat channels like Visa, WeChat Pay, Alipay, and UnionPay alongside high-efficiency crypto settlement networks—enabling builders worldwide to secure the compute resources they need with minimal transaction friction.
In an era where tokens are currency, B.AI is dedicated to serving as the foundational compute layer for AI innovators, continuously strengthening its super-compute hub to power thousands of industries. Here, accessible compute is no longer just a narrative—it is tangible infrastructure driving productivity and intelligent transformation for enterprises worldwide.
B.AI Team
support@b.ai
Disclaimer: The views, suggestions, and opinions expressed here are the sole responsibility of the experts. No Economy Extra journalist was involved in the writing and production of this article.
