❌

Normal view

DeepSeek is hiring 150 engineers, and none of them will touch a model

abstract bubbles

Hundreds of thousands of AI agent sandboxes can already run concurrently on a single DeepSeek cluster. Now the company is staffing up to handle what happens as that number — along with its training, evaluation, and other backend workloads — keeps climbing.

Cui Tianyi, who joined DeepSeek in March and works on its Harness team, the group responsible for the infrastructure and environments used to run and evaluate agents, announced in an X post that roughly 150 engineering positions on September 7, with the hiring concentrated in server-side engineering and Agent Elastic Compute rather than AI research. The work spans operating systems, virtualization, networking, storage, scheduling, and the control-plane services that coordinate those resources.

Cui said DeepSeek’s existing backend systems will need upgrades, maintenance, and rewrites as workloads grow. One such system at the center of that scaling challenge is DeepSeek Elastic Compute, or DSec, the sandbox infrastructure DeepSeek built to execute agent workloads during post-training and evaluation.

Cui said DeepSeek’s existing backend systems will need upgrades, maintenance, and rewrites as workloads grow.

Four sandboxes, one SDK

Agent workloads require more than GPUs for inference, with each agent also needing an isolated environment to run code, call tools, change files, and collect the results.

DSec supports four types of those environments through the same Python SDK. Simple function calls go to pre-warmed containers, while Docker-compatible containers handle jobs that need a persistent environment. DeepSeek uses Firecracker microVMs when stronger isolation is needed and QEMU virtual machines for workloads that require a full guest operating system.

That range means the same infrastructure can handle anything from a simple tool call to a software-engineering task that needs an entire OS. It’s a similar challenge to the one the rest of the industry is bumping into as agents move from demos to production. OpenAI, for instance, recently designed custom silicon specifically to address the compute pressure that agent workloads create, and DeepSeek open sourced its own agent harness in August.

Lazy loading agent environments

Every sandbox needs its own environment, but copying complete container or VM images onto every host would consume enormous amounts of storage and network bandwidth while adding to startup time. DeepSeek gets around that by tying DSec into 3FS, the distributed filesystem it originally built for its AI infrastructure, and keeping container base images and filesystem commits as read-only layers backed by 3FS.

The metadata stays local, but the underlying data blocks are fetched only when they’re actually needed. MicroVMs use a similar setup, sharing their read-only base layer through 3FS while writes from individual sandboxes are kept in local copy-on-write layers.

DeepSeek says DSec reduces duplicate page-cache usage across virtualized environments and reclaims memory to allow safe overcommitment, while changes to the container runtime cut the CPU overhead of each sandbox.

The team also had to deal with spinlock contention inside the container runtime. At small scale, the CPU time spent there barely registers. At scale, it limits how densely those environments can be packed onto each host.

DeepSeek says DSec reduces duplicate page-cache usage across virtualized environments and reclaims memory to allow safe overcommitment, while changes to the container runtime cut the CPU overhead of each sandbox.

When replay breaks training

During reinforcement learning and other post-training workloads, large numbers of agent rollouts can be running at once, and jobs may be interrupted as compute gets reassigned. Starting over wastes everything the agent has already done, but picking up where it left off isn’t as simple as replaying its previous commands.

Some of those commands may have changed a file or otherwise altered the environment, so running them again could produce a different result or leave the training trajectory in the wrong state. DSec avoids that with a globally ordered trajectory log that records commands along with their results.

When a rollout resumes, DSec can fast-forward through the completed work using those recorded results rather than executing the commands a second time. That reduces the cost of interruptions across thousands of training and evaluation runs, while the same logs preserve a history of how each sandbox changed and allow earlier sessions to be replayed.

Engineers, not researchers, wanted

The roughly 150 openings reach across DeepSeek’s backend, including the lower-level systems work behind Agent Elastic Compute as well as the services that support its models and agents.

DeepSeek said in June that it planned to at least double the size of every department, but this round of hiring leans heavily toward the systems underneath its models rather than the models themselves. DSec is part of that work, with hundreds of thousands of sandboxes running concurrently and putting pressure on everything from how jobs are scheduled to how they recover after an interruption.

The roughly 150 openings reach across DeepSeek’s backend, including the lower-level systems work behind Agent Elastic Compute as well as the services that support its models and agents.

The post DeepSeek is hiring 150 engineers, and none of them will touch a model appeared first on The New Stack.

Could GLP-1 Drugs Help You Live a Longer, Healthier Life?

8 September 2026 at 17:33

Elderly mice on the popular weight-loss drug semaglutide aged slower and lived longer, mimicking the longevity effects of caloric restriction.

It’s a hot GLP-1 drug summer. Semaglutide—better known as Ozempic and Wegovy—seems to be everywhere. The blockbuster weight loss aid mimics a natural hormone that tells the brain “you’re full,” making it easier to shed pounds without the constant hunger and cravings that make traditional dieting miserable.

The drug may also have longer effects. A new study in mice suggests that starting semaglutide in old age extends lifespan roughly 12 percent. The mice regained their curiosity and memory and improved on myriad age-related hallmarks. Chronic inflammation cooled. Senescent “zombie cells” dwindled. Their genomes and proteins became more stable, and the hippocampus—a brain region crucial for learning and memory—sprouted new neurons.

The findings are reminiscent of calorie restriction, a way to boost longevity and stave off age-related health problems, at least in lab animals. But sticking to a diet for years, let alone decades, is tough. Scientists have long searched for a drug that could deliver some of the benefits without hunger pangs. Semaglutide seems to fit the bill, with surprising perks beyond dieting.

“If these results from mice hold true in people, this is a promising hint that people taking GLP-1R agonists [GLP-1 drugs] may have an age-slowing benefit as well as benefits to reducing diabetes and obesity,” said Tara Spires-Jones at the University of Edinburgh, who was not involved in the study.

But don’t go ordering GLP-1 drugs online just yet. The study used inbred female mice, whose physiology differs from that of women aging through or beyond menopause. Rapid weight loss in people taking these drugs can also reduce lean muscle, potentially increasing fragility in older people. And while GLP-1 drugs are already being tested in the battle against neurodegenerative disorders, whether they sharpen the aging brain remains an open question.

Still, longevity researchers are cautiously optimistic.

“The findings reframe a key question that has often been asked back to front,” wrote Maria Fernandez and Rafael de Cabo at the National Institute of Aging, who were not involved in the study. “Rather than considering the broad health benefits of GLP-1 drugs as something to be explained one disease at a time, these results suggest that the positive effects have a common cause: slowed aging.”

Fountain of Youth

One way to live healthier and longer is seemingly mundane: adopt a healthy lifestyle.

Diet and exercise have repeatedly been linked to better health in our twilight years. As we age, our bodies slowly break down. Damage and mutations accumulate in DNA. Telomeres, the protective caps at the ends of chromosomes, grow shorter, contributing to genomic instability. The molecular switches that turn genes on or off go haywire, and our cells become less able to make working proteins and clean up damaged ones.

There’s more. Mitochondria, the cell’s power plants, struggle to produce energy and leak toxic molecules. Damaged cells stop dividing, lose their function, but stubbornly refuse to die. Instead, they leak a toxic chemical soup that damages tissues. Chronic inflammation flares, and stem cells run out of steam, making it harder for the body to regenerate or repair itself.

Together called the hallmarks of aging, this laundry list has long challenged scientists looking for a silver bullet against the march of time. They have discovered some exotic options. Transferring components of young mice’s blood to older recipients has rejuvenated faltering hearts, kidneys, and brains. And clearing senescent “zombie” cells with drug cocktails or genetic engineering has attracted billions of dollars in investment.

But perhaps the most studied, and most robust, intervention is caloric restriction. In flies, rodents, and other lab animals, slashing energy intake without causing malnutrition has repeatedly extended lifespan and delayed multiple age-related diseases. Restriction triggers broad changes, including improved metabolism and insulin sensitivity. It also helps prevent damage to cells. In humans, moderately reducing calories seems to improve heart health and slow the speed of biological aging.

Sticking to a diet for years on end, however, is hardly sustainable.

Cheat Code

GLP-1 drugs may make it easier.

Originally designed to control blood sugar and appetite, the drugs have also shown promise for reducing rates of cardiovascular, kidney, liver, and neurodegenerative diseases, at least in people with obesity or Type 2 diabetes. Clinical trials are now exploring their effects in people with metabolic liver disease, which becomes more common and consequential with age.

Not all these effects can be explained solely by weight loss, raising a bigger question: How can a single class of drugs influence so many seemingly unrelated conditions that often crop up with age?

“If GLP-1 medications slow down the aging process itself, a wide range of clinical benefits is

exactly what would be expected, because aging is the root of most chronic diseases,” wrote Fernandez and de Cabo.

To test that theory, the team gave daily semaglutide injections to 20-month-old female mice—roughly comparable to women in their early-to-mid 60s—for as long as they lived.

Compared to a group of mice given saline, the treated mice lived an average of 834 days, versus 724 days for a control group. Both groups could feast on standard chow to their hearts’ content. After just three months of treatment, the mice taking semaglutide were more lively and curious than their peers. They readily explored new environments, balanced better on a skinny rotating rod, ran faster on a tiny treadmill, and solved mazes more quickly.

Under the hood, semaglutide blunted the hallmarks of aging across the board. Stem cells in the bone marrow and hippocampus sprouted, suggesting renewed regenerative capability. DNA damage and protein and energy dysfunction declined. Zombie cells partially disappeared.

The results weren’t simply a consequence of eating less. In another three-month experiment, the team compared the drug with a calorie-restricted diet that cut energy intake by 24 percent—the same reduction seen in the semaglutide-treated mice.

The drug seemed to have a leg up. Compared with caloric restriction, it produced more improvement in cognition and blood sugar control, without the metabolic adaptations normally driven by hunger, such as lowered energy expenditure. The mice also showed fewer signs of hunger. They ate on a normal schedule rather than prowling for food before feeding time and gobbling rations once available.

Genetic sequencing of the liver found semaglutide triggered similar molecular signaling pathways as caloric restriction, like for example, those that sense nutrient availability, cell stress, and proteins involved in longevity. Both interventions tamped down inflammation and boosted genes involved in handling fats, but they didn’t follow the exact same biological playbook.

The findings raise the “intriguing question of whether GLP-1 drugs target an alternative biological route into aging that has its own side effects and therapeutic ceiling,” wrote Fernandez and de Cabo. Exactly how that route works remains unclear, but the team is eager to find out.

The findings are promising, but the study also has major limitations. It didn’t directly compare dieting and semaglutide for lifespan extension. And because gender affects the aging process, the team will need to see if the results hold in males. Then there’s the potential loss of lean muscle, a serious concern for people who are already frail or saddled with age-related health conditions.

Even so, “semaglutide is probably the best caloric-restriction mimetic I have seen,” Tim Rhoads at the University of Wisconsin–Madison , who was not involved in the study, told Chemical and Engineering News.

Untangling semaglutide’s bonus effects could ultimately reveal new ways to slow—or even rewind—aging’s ticking clock.

The post Could GLP-1 Drugs Help You Live a Longer, Healthier Life? appeared first on SingularityHub.

Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses

8 September 2026 at 14:15

Avid Artifacts readers know that we have been covering not only models but also their licenses for quite some time. There was a period when custom licenses were all the rage, for example the custom Qwen2.5 72B-Instruct license or the Llama licenses. DeepSeek had a custom license for DeepSeek V3 before R1 changed it to MIT, which has resulted in many (Chinese) model makers adopting MIT or Apache 2.0 licenses in 2025.

In 2026, open models are more competitive than ever, which has led to two interesting developments: Western model makers adopt open licenses, with both Google and Meta switching to Apache 2.0. Chinese model makers at the frontier, however, are becoming more restrictive: Kimi K3 comes with a license which requires commercial agreements for those who run inference or fine-tuning services, and MiniMax M3 requires agreements above a revenue threshold and has prohibited use cases.

The newest addition is Zhipu’s GLM-5.3, which switched from MIT (GLM-5.2 and earlier) to a custom license with the following clause for inference and fine-tuning providers:

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose. The scope and method of the security review shall be reasonably determined by Z.AI.

While the 10 billion US dollar threshold is very high compared to other licenses of this kind, “affiliates” is not defined in the license, which adds uncertainty and creates barriers to adoption. Furthermore, the license is provided in both English and Chinese, with the Chinese text using “关联方” for affiliated parties, which does have a definition in Chinese law.

We are by no means legal experts and there are obvious reasons why those licenses are created. However, we want to highlight the issues that come with creating such licenses, especially in a world with a lot of valid open and closed alternatives.

Share

Our Picks

  • Motif-3 by Motif-Technologies: Motif is one of the few hidden gems out there, showcasing innovation in their model training with very limited resources compared to others. Motif-3 comes with an MIT license and impressive scores for its size. Given the trajectory of model releases from Motif 2.6B, which we covered in 2025 and Motif-2-12.7B, the improvements are impressive.

  • dots3-note-prev by dots-studio: RedNote/Xiaohongshu, the Chinese Instagram, is also getting more serious about model training, although they aren’t exactly a newcomer, having released models as early as 2025. dots3 was also able to win the IMO 2026 with a perfect score using an internal harness. We expect more from them in the near future.

    General Reasoning and Agent evaluation results
  • Qwen3.8-Flash-Next by Qwen: A preview of the next version of Qwen models in terms of architecture: 125B-A6B with 51B n-gram embeddings. It uses GDN and Qwen Sparse Attention. Similar to Qwen3-Next-80B-A3B-Instruct, we expect similar architectures to become more popular and the ecosystem to fix integrations by the time Qwen4 drops.

  • GLM-5.3-Flash by zai-org: This release perfected the version of the Chinese model playbook we’ve written about in 2025: The model got released as a free-to-use “stealth model” under the name “Ox-Alpha” on OpenRouter and OpenCode, which got people excited to try it out in the first place. They then speculated about its creator and size, alleging it is a >1T model from Cursor/xAI, Gemini or a new pre-train from open source labs. Because the model is relatively performant, people kept speculating for days about its creator, thus building up hype. It also dampens the accusations of benchmaxxing which accompany every (open) model release.

    bench_53
  • Hy4-preview by tencent: Tencent is becoming a serious player in the open model space, increasing the size of their flagship model while spinning the post-training flywheel. The result, Hy4-preview, is a competent model which currently has an issue with overthinking. However, if the trajectory from Hy3-preview to Hy3 is any indication, the final model might be a legit shot at the front ranks of open models.

View more details on all the models in this issue at our Artifacts Hub.

Visit artifactshub.ai

Models

General Purpose

  • NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 by nvidia: An update to Nemotron, which comes with performance — but especially speed improvements — across the board.

    bench_53
  • Ling-3.0-flash by inclusionAI: Ant Ling is a frequent guest at the Artifacts Log; they are now on their third iteration of models, adopting a hybrid design (KDA + Gated MLA), similar to others. They also release a small 7.9B-A1.3B version.

  • Qwen3.8-2.4T-A95B by Qwen: In a rather surprising turn of events, Alibaba started to openly release their biggest versions of Qwen as well. However, it comes with a custom license and its performance is behind other models of its size.

Read more

After nine years as HashiCorp CEO, Dave McJannet now wants to “unblock” enterprise AI agents

Retro 3D-rendered computer with a two-icon logo on screen, keyboard, and mouse on a purple background

Ask a traditional enterprise application for a customer address or today’s revenue figures and, broadly speaking, it follows a predictable route its developers have already mapped out: authenticate the user, query the right system, return the result. Given the same underlying data, you’ll get the same answer each time.

Ask an AI agent the same question, and the journey is much harder to forecast. It might consult one system, decide it needs more context from another, make a dozen tool calls, pass information through a language model and only then produce an answer. Run the same request again, and it may take a different route altogether.

And in an enterprise, what happens along that route can matter just as much as the answer: which systems the agent accesses, what data it sees, what actions it takes and how much it spends.

That distinction — between predetermined software, and applications that make probabilistic decisions on the fly — sits at the heart of a new company from a founder who knows a thing or two about bringing order to a new generation of infrastructure.

AI agents are hard to govern

Dome Systems co-founder David McJannet left HashiCop in August 2025
Dome Systems co-founder David McJannet left HashiCop in August 2025

Dome Systems was co-founded at the turn of the year by David McJannet, who spent close to a decade leading Terraform-creator HashiCorp through the cloud era, culminating in its blockbuster 2021 IPO and subsequent $6.4 billion sale to IBM in 2025. McJannet is joined at the helm by Marc Holmes, who spent more than six years at HashiCorp as chief marketing officer.

In an interview with The New Stack, McJannet lays out his company’s thesis on AI agent governance, arguing that enterprises are now running into the same kind of problem that they did with cloud infrastructure: adoption comes first, then the real spadework begins of putting the right controls in place across security, operations and finance.

“It’s actually a very different architecture, and that is what unlocks the power of these new [agentic] applications.”

Part of the challenge, he says, is that agents are built very differently from the enterprise applications of yore, which companies spent years learning how to control.

“It’s actually a very different architecture, and that is what unlocks the power of these new [agentic] applications,” McJannet explains.

He points to self-driving cars as an example: a model takes in live inputs and interacts with the vehicle’s systems as conditions change, because no developer can reasonably pre-program every possible situation a car might encounter on the road.

“It’s making judgments along the way, as opposed to trying to look up the historical maps of the world and make a real-time decision,” McJannet continues.

An enterprise agent can behave in much the same way: call one tool, assess the result, decide it needs another, and keep going until the task is complete. That flexibility lets agents tackle work that would be difficult to script exhaustively in advance — but it also makes their behaviour harder for enterprises to govern.

And this gets to the heart of what McJannet is striving for with Dome.

Table stakes for the agent era

The company launched out of stealth back in April with $14 million in seed funding, with McJannet having departed HashiCorp the previous August after the IBM transition concluded.

Dome’s starting point is that an agent combines three things: code, a model, and the backend systems or tools it interacts with. Bringing those pieces together under one platform, McJannet says, is “table stakes” for applying meaningful constraints to what the agent can do.

“If you don’t have an integrated platform, you can’t enforce controls across everything that the agent is doing,” McJannet says.

“If you don’t have an integrated platform, you can’t enforce controls across everything that the agent is doing.”

And so Dome’s platform is built around those three elements. An agent registry keeps track of the agents themselves; an MCP gateway controls the tools they can call; and a model broker/router governs which models they can use and how requests are routed.

The setup starts by registering the agent and giving it an identity, establishing who is allowed to call it, and connecting the backend tools it can reach — Zendesk, in this example.

Dome registers an agent, verifies its caller and connects the tools it can use.
Dome registers an agent, verifies its caller and connects the tools it can use.

Next, Dome connects a model provider, groups available models into a pool with routing and failover rules, then combines the agent, its tools and its models behind a single gateway. That gateway becomes the point through which Dome can apply the policies governing what the agent is allowed to do.

Dome connects a model provider, creates a model pool and brings the agent behind a gateway.
Dome connects a model provider, creates a model pool and brings the agent behind a gateway.

Once those pieces are connected, teams can set permissions on each call, use guards to inspect responses, apply quotas to cap spending, and keep a common audit trail across the agent’s activity.

Today, McJannet says, enterprises are often piecing all of this together themselves. A standalone model broker might be brought in to control spending, while a separate tool gateway handles security and operational concerns. Some are then building their own agent registry to tie those systems together.

Moreover, buying those capabilities separately leaves enterprises with another integration problem to solve. A model router might govern one part of an agent’s activity and a tool gateway another, while the agent itself continues moving between them.

“If you just provide the tool gateway or just the model router, it doesn’t allow you to have this kind of system of control,” he says.

That is also where Dome’s latest move enters the fray. After spending its first months in early access, the company is now opening the platform to self-service users for the first time, allowing teams to sign up with little more than a credit card, bypassing the typically arduous enterprise sales process.

Dome goes self-serve

Self-serve is relatively unusual route for this kind of enterprise infrastructure product. Dome is publishing its prices, offering a free tier and letting practitioners get started without first going through a sales process, while keeping the traditional enterprise route open for larger customers.

The thinking is partly about who McJannet expects to use the product. Rather than limiting access to buyers who are already deep into a procurement process, for example, self-serve enables individual practitioners to be able to discover, try and use the platform themselves.

“”We want to make the barrier as low as possible to have people come on board,” McJannet says, adding that Dome had already seen a number of self-service sign-ups ahead of the launch.

Separately, its pricing reflects a belief about where value will ultimately sit in this market. McJannet regards model routing and tool connectivity as baseline capabilities, with the more valuable piece being the controls that sit across the agent as a whole — think permissions, data redaction and spending quotas.

It’s also worth noting that while Dome’s main target user will be platform engineering teams inside large enterprises, typically working alongside operations and security, self-serve also creates an opening for another kind of user: the small company, perhaps even only one or two people, building an agent and trying to sell into an enterprise. The sort of scenario that aligns with the fabled one-person unicorn promised by many in the AI realm.

Indeed, McJannet says developers can get far building the application itself, only to hit a wall when a prospective enterprise customer begins its security and operations review. How is identity enforced? Who can see the data the agent reaches? What happens when it calls other agents? Can its activity be reconstructed afterwards?

Some builders, he says, have asked whether they can “certify” their agents on Dome because “my agent won’t get deployed until I can satisfy these infrastructure elements.” McJannet is careful to add that Dome doesn’t currently run such a certification program, but it’s clearly one route the company could venture down.

“If you register that agent on Dome, all the infrastructure elements are taken care of,” McJannet says.

‘Unblocking AI agents’: Lessons from the cloud era

That division between developers eager to ship, and enterprise teams worried about what happens after, is also where McJannet sees the strongest parallel with his years at HashiCorp.

During McJannet’s tenure, HashiCorp increasingly positioned itself around helping large organizations standardize how cloud infrastructure was provisioned, secured and connected. That included the 2020 launch of HashiCorp Cloud Platform (HCP), which offered its infrastructure tools as managed cloud services.

More broadly, McJannet’s account of early cloud adoption begins with developers swiping a credit card and deploying directly to Amazon because cloud infrastructure allowed them to build applications that had previously been impractical. The applications were compelling enough that enterprises adopted cloud despite resistance from operations and security teams, and what followed was a second phase: companies needed common services for provisioning, credentials, networking and other controls before cloud could become routine across the organization.

Platform engineering teams became the people responsible for reconciling those two demands: allowing developers to build while giving security, operations and finance enough control to permit those applications into production. McJannet believes agents are now creating the same tension.

“You’ve got this queue of cool apps that developers build that the ops and security teams are just not comfortable letting flourish in their environments.”

“You’ve got this queue of cool apps that developers build that the ops and security teams are just not comfortable letting flourish in their environments,” he says. “And so, inevitably, it has to go that same direction where the platform engineering team has to figure out [a way] to get to say ‘yes’.”

Dome’s bet is that enterprises will eventually prefer one system spanning the entire agent to a patchwork of gateways, routers and security products. In McJannet’s telling, that common control layer is what gives enterprises a way to limit how far an agent can roam while still letting it act autonomously.

“You have to have this control layer that provides this corridor where we can constrain the behavior of that new type of application architecture,” he says. “Because without that, you cannot unblock the deployment of AI applications.”

“That’s the part that we’re trying to answer — how do we unblock agents at scale?”

There is still plenty for Dome to prove. The company isn’t naming customers at this stage; McJannet says none of the enterprises it has worked with are yet willing to be identified publicly, though he says Dome has spent the past eight months talking to dozens of them.

Ultimately, McJannet believes the cloud era showed that new applications only become commonplace once enterprises have the controls to let them through. Dome is his attempt to solve that problem for agents.

“I think that’s the part that we’re trying to answer — how do we unblock agents at scale?”

The post After nine years as HashiCorp CEO, Dave McJannet now wants to “unblock” enterprise AI agents appeared first on The New Stack.

Introducing CUDA Rust: Two Tracks for Writing GPU Kernels

8 September 2026 at 12:00
In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and...

In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and NVIDIA will be growing and maturing CUDA Rust into 2027 and beyond The systems layer of AI spans inference engines, serving infrastructure, drivers, and agent runtimes, and it churns constantly as models and techniques change.

Source

How llm-d makes the most of the hardware you already have

8 September 2026 at 12:00
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
❌