DeepSeek Unveils DSec Sandbox Infrastructure for Large-Scale Agent Training

DeepSeek's DSec platform scales agent training with millions of sandboxes, reshaping AI development.

Last Updated: September 23, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: September 23, 2026

September 23, 2026, (Inside AI) — DeepSeek has unveiled details of DeepSeek Elastic Compute, or DSec, a sandbox infrastructure built to train AI agents at a scale previously unseen. A paper posted to arXiv and authored by more than 130 researchers, including founder Liang Wenfeng, describes a system that unifies four types of sandboxes and coordinates them with reinforcement learning workloads.

The paper reports that a production-scale DSec unit spans roughly 160 nodes, supports about 3 million sandboxes per day, and handles more than 380,000 concurrent sandboxes. It can create over 5,000 sandboxes per second. These figures suggest a shift from experimental agent training to industrial-grade infrastructure.

Sandboxes are isolated environments where AI agents can execute code, call functions, and interact with simulated worlds without risking real systems. DSec combines function-call, container, microVM, and full-VM sandboxes into a single platform. This unification allows developers to choose the right isolation level for each task, from lightweight function calls to full virtual machines.

The system's tight integration with reinforcement learning is key. In RL, an agent learns by receiving rewards for actions. DSec coordinates the lifecycle of each sandbox with the training loop, spinning up environments on demand and tearing them down when episodes end. This reduces idle time and increases throughput.

Read: DeepSeek Launches Harness Framework for Autonomous AI Agents

DeepSeek's move reflects a broader industry trend. As AI agents become more capable, the need for scalable, secure testing grounds grows. Competitors like OpenAI and Anthropic have developed similar sandboxing tools, but DeepSeek's published metrics emphasize raw scale. The involvement of Liang Wenfeng, who rarely appears in technical papers, signals the project's strategic importance.

"The ability to run millions of sandboxes daily is a game-changer for agent training," said a researcher familiar with the paper, who spoke on condition of anonymity. "It means you can parallelize exploration at a level that was previously impractical."

Inside AI could not independently verify the performance claims. However, the paper's detailed architecture and the reputation of its authors lend credibility. The system's design also addresses a common bottleneck: the overhead of managing many short-lived environments. By automating lifecycle management, DSec frees researchers to focus on algorithm design.

While the paper focuses on training, the implications extend to deployment. Agents trained in diverse sandboxes may generalize better to real-world tasks. This could accelerate the adoption of AI in fields like software development, where agents must navigate complex codebases.

DeepSeek has not announced a commercial product based on DSec. The paper's release suggests a commitment to open research, but the company may keep the infrastructure proprietary. For now, the AI community gains a blueprint for building high-scale agent training systems.

Read: Baidu Launches DuMateBench Benchmark for Real-World AI Agent Delivery

As agentic AI moves from demos to production, infrastructure like DSec will be critical. The ability to train agents at scale could determine which companies lead the next wave of AI innovation. DeepSeek's paper provides a rare look at the engineering behind that race.

More from Inside AI

  • AI In Business

    How AI Startups Can Attract Funding and Talent in Their First Year

    September 23, 2026
  • AI Policy & Regulation

    OpenAI, Anthropic CEOs to Brief UN Security Council on AI Risks

    September 23, 2026
  • Agentic AI

    Qualcomm Exec: AI Agents Will Make Apps a Second Option

    September 23, 2026
  • AI Hardware & Infrastructure

    Alibaba Unveils Zhenwu V900 Chip, Plans 10 Trillion Parameter AI Model

    September 23, 2026
  • Agentic AI

    UPSC Mains 2026: Agentic AI in Digital Payments and Climate Risks to India’s Economy

    September 23, 2026
  • AI Tools

    AWS Launches Amazon CloudWatch Omni for AI Agent Observability

    September 23, 2026
  • AI Policy & Regulation

    UK Ready to Lead Global AI Standards Push, Burnham Tells UN

    September 23, 2026
  • AI Safety

    Ex-Google Safety Chief Warns AI Could Harm Children More Than Social Media

    September 23, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Join Our Newsletter Community

Subscribe

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content
  • Advertise with us
  • Newsletter

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital