Last October, the P99 conference — the online gathering for developers focused on high-performance, low-latency applications — featured a cracking keynote from Chip Huyen.
The author of the best-selling AI Engineering, Huyen opened with simple math: Training a frontier model is a one-off cost, but inference is the same cost paid over and over. That’s great for the frontier model providers, and bad for us token burners. Over the life of a model, Huyen reckons the compute split ratio lands somewhere between 1:10 and 1:100 for training to inference. Reasoning models — which burn even more tokens — push that out even further. We all know the feeling of hitting our weekly session quotas.
Huyen’s point is that if inference is too expensive, then nobody ever recovers the training bill, which might explain why there are so many memes about the “profitability” of frontier models. So, how do we optimize inference?
That’s a topic that Huyen spent months researching for her book. And in the spirit of optimization, Huyen distilled it down to 30 minutes for the conference in October 2025.
Huyen is returning for P99 CONF 2026 in a few weeks. Ahead of that moment, watch her full talk – or read the recap below – from last year, then let’s talk about how those ideas aged over the past 11 months.
What to measure
Chip recommends focusing on a few key latency metrics:
Time to first token (TTFT): How much time elapses before the user sees anything
Time per output token (TPOT): The average time between consecutive tokens (aka inter-token latency)
End-to-end latency: Time to first token, plus time per output token, multiplied by the number of output tokens minus one
(Click to enlarge graphic.)
With reasoning models, some of those tokens never reach the user. “The first generated token might not be the same as the first visible token,” Huyen explained. “The model might think for a while, and it will only show the first token of the final output to the user.”
“The first generated token might not be the same as the first visible token,” — Chip Huyen
Some people also measure Time to Publish for that (i.e., how long until the user sees the first token). The best metric to prioritize depends on what matters most for your users.
Also consider “goodput” alongside throughput. Throughput measures requests processed in a given window. Goodput measures the requests that actually met your targets. Chip’s example: an app targets 200 ms time to first token and 100 ms time per output token, and processes 10 requests per minute, but only three hit both.
(Click to enlarge graphic.)
3 ways to optimize LLM inference
With inference servers, you can optimize from 3 different angles: the hardware, the model, and the service that manages the requests and responses.
(Click to enlarge graphic.)
Huyen previously worked at Nvidia and opted out of the hardware discussion: “Even though I find it to be an intellectually interesting topic, it’s not relevant to a lot of people because we don’t have the power to change the hardware itself,” Huyen explained. She also didn’t want to spend much time on the obvious solution: replica parallelism, or just adding more machines. It’s costly, and it gets complicated fast – especially if you end up with a mix of 80GB, 48GB and 24GB machines and models of varying sizes to distribute across them.
That leaves the model and the service. Huyen offers these tips on how to decide: “If you want to host the models yourself, or if you have access to the model weights, or if you train a model yourself, or you want to fine-tune or distill a model, then model optimizations might be for you. However, if you want to take a model as-is and make it more efficient on your own inference service, you might want to look into service optimizations.
Model optimization
The following techniques change the actual weights so that they can change the model outputs.
Quantization lowers the precision used to store weights and activations (e.g., from four bytes per parameter at 32-bit to one byte at 8-bit). Huyen explained, “Reducing the precision not only reduces the memory requirement to run the model, making it cheaper. It can also make the model a lot faster. If you do additions bit by bit and each weight is 32 bits, you have to do it 32 times. If it’s 8 bits, you only have to do it eight times.”
The tradeoff is a small quality hit. Huyen continued: “It’s possible to reduce a lot of the model’s memory footprint with minimal quality degradation, and quantization is pretty generalizable to a wide variety of model architectures and model sizes. That’s why it’s very popular. I rarely see any companies running a model at full precision anymore.”
“I rarely see any companies running a model at full precision anymore.” — Chip Huyen
Distillation involves using a large model to generate training data for a smaller model. For example, say you have a truly large model (the example Huyen used was o1) and want a model that performs like it, but is much smaller. Basically, you collect a large set of prompts, run them through the larger model, then train the smaller model on its responses.
Proceed with caution, though. Huyen warned, “A lot of model providers have the condition that they do not allow their models to be used to train competitive models. So even though it’s a very common technique, you need to check licensing.”
Service optimization
This set of techniques targets how requests are scheduled, routed, and reused. The actual weights aren’t affected.
Batching groups multiple requests so they’re processed together in a single pass through the model – which is much more efficient than dealing with them one at a time. Huyen presented a few batching options:
Static batching waits for the batch to fill. This maximizes compute utilization, but it might increase the latency for the first requests.
Dynamic batching runs on a timer instead (e.g., batching every 15 ms). This is less compute-efficient, but it’s better for latency.
Continuous batching handles the case where requests finish at wildly different times, which is common with LLMs. One request asks for the capital of Vietnam; another kicks off deep research. With static or dynamic batching, the finished request’s slot sits idle until the slowest one completes – and new requests queue up behind it. Continuous batching returns each request as it finishes and fills the spot with another request. That can improve compute resource utilization and latency.
(Click to enlarge graphic.)
Decoupling prefill and decode separates the two phases of a request onto different machines. (Prefill processes the input, while decode generates the output.) Huyen said, “Input tokens can be processed in parallel, whereas output tokens need to be generated sequentially. With parallel processing, it’s bounded by compute, the processing power of the chip. With decoding, it’s bounded by memory, because you have to move model weights.”
Because each phase stresses different resources, most services now separate them. To improve time to first token, shift machines toward prefill. If you care more about improving time per output token, shift them to decode.
(Click to enlarge graphic.)
Parallelism splits work across machines. Replica parallelism copies the whole model onto more machines. Tensor parallelism divides a very large matrix, so different machines compute different parts of it. Pipeline parallelism divides the model by layer, so requests move through as a pipeline.
(Click to enlarge graphic.)
Prompt caching processes shared text once, saving cost and latency. A lot of repetition exists across requests to the same application: the system prompt, the examples, the same code base, the same document behind different questions. You might as well process that shared segment once, cache it, and reuse it.
The technique was relatively rare when Huyen was writing AI Engineering. “There was one paper about it, and it was not really known, but it made a lot of sense. So I included prompt caching in the book, and I’m very happy to see that nowadays it’s pretty much everywhere.”
(Click to enlarge graphic.)
The savings scale depending on how much of your prompt gets cached. In Claude Code logs, Huyen’s open-source tool Sniffly found cache hit rates of 90% to 97%. Some providers rewrite prompts internally to improve hit rates, but you might as well structure them yourself.
Huyen’s tip: since caching works on shared prefixes, put the stable parts of your prompt first and the variable parts later. “It’s pretty easy to do, and it can improve your application performance significantly,” she noted.
Evaluating inference providers
Huyen closed with a warning for anyone evaluating inference providers: “There are many inference companies that provide inference optimizations for models you want to use, and a lot of them advertise just cost and latency.
“But pay attention to how many inference optimization techniques also change the model behavior or reduce the model quality. So when evaluating an inference service, it’s important to look not just at cost and latency, but also at model quality. Does this model, provided on this service, also perform similarly on standard benchmarks?”
What’s changed one year later?
So where do we stand today, one year on from this keynote? Most of it actually aged quite well.
On the economics, I reckon Huyen was bang on… I think, for most of us as users, we don’t have all the cost levers to pull that Huyen outlined. But it’s great to understand what is happening. As a novice local LLM user myself, I found I could relate to her points on parallelism (I don’t have it) and prompt caching/quantization (within my grasp of control).
Prompt caching (which Huyen said was new when Huyen wrote AI Engineering) is now priced into every bundle purchase of API tokens. And her Claude Code observation (90% cache hit rates) is probably the reason we mere mortals can still afford agentic coding agents at all.
Some of it aged in ways that were hard to predict at the time. Huyen mentioned how reasoning models make inference even more significant. One year on, I think agents running multi-step loops with tool calls have turned that idea from a footnote into a way to turn Claude’s rate limits (and their infamous 99.x% availability) on their head.
All the metrics Huyen described – time to first token, time to publish, goodput under a latency SLO, etc., are all now part of the lingo and probably need to be reasoned about differently.
That’s one thing I hope she’s talking about this year! Grab a free conference pass and join us online.
Grab a complimentary pass to PG 99 Conf 2026 and join us on October 21 and 22 to chat with Huyen.
Enterprise AI company Cohereannounced North Small Translate last week, a mixture-of-experts (MOE) open-weight machine translation model that works across 50 languages.
Developers can download the weights for noncommercial use under CC BY-NC 4.0. Cohere offers commercially licensed deployment through Model Vault, which is a Cohere-managed inference environment. Cohere positions the model as part of its sovereign AI strategy, aimed at organizations that want greater control over where their models run and how their data is handled.
North Small Translate builds on Cohere’s multilingual and translation lineage, which includes its Tiny Aya and Command A Translate model families. The company claims North Small Translate outperforms “similarly sized open-weight models” under 1T parameters, as well as API-based translation models in various dimensions of machine translation on average.
Cohere co-founder Nick Frosst tells The New Stack that the model’s efficiency draws from the fact that it is non-reasoning, i.e., it relies on learned statistical patterns without a step-by-step logic process, which means it uses fewer tokens.
Machine translation is still broken for most of the world’s languages
“We spent nine years scaling an architecture invented to fix translation, and machine translation is still broken for most of the world’s languages,” Frosst says. “General-purpose models get you most of the way and then stop. The next phase of enterprise AI in this space is smaller, more specialized, and runs inside your own walls.”
“…machine translation is still broken for most of the world’s languages.”
In Cohere’s reported evaluation using WMT26 benchmarks, the company states that North Small Translate leads with a WMT26 All Languages benchmark score of 83.60, compared with 81.56 for Qwen 3.5 397B A17B, 76.50 for GLM 5.2 FP8, 81.37 for DeepL NextGen, 79.46 for Gemma 4 31B (on), and 68.20 for Google Translate.
With its mixture-of-experts architecture and 218 billion total parameters, with 25 billion active. Cohere points to North Small Translate’s smaller compute & memory footprint than other models. Some model-to-model comparisons in this space aren’t fully substantiable, since not every vendor discloses parameter counts.
With current solutions, long documents start to fall apart
“Machine translation allows documents to be translated from one language to another automatically. With current solutions, long documents start to fall apart,” Frosst says. “Google Translate scores 21.3 on our long-context test, Gemma 4 31B 19.4; we score 48.9. That’s [for example] a safety manual that reads fine on page one… and has drifted by page ten. The other risk is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building, and necessarily that means your control over it is diminished.”
“The risk [in machine translation] is where the text goes. Once you push HR policies or regulated documents through a third-party API, that data has left your building and necessarily that means your control over it is diminished.”
Explaining why the model offers “stronger translation performance” across complex enterprise translation tasks, Frosst says the model can support work spanning “a high volume” of sensitive documents.
As well as its 50 languages (32 ‘high-resource’ languages + 18 others), the Cohere team explains that the model also supports translation-workflow-focused capabilities, such as structured translations (i.e., Markdown or JSON documents), instruction following (i.e., recommended tone & format), and terminology guides (i.e., providing specific vocabulary to use in the translation), all as part of the model.
“North Small Translate works with a multi-pass workflow,” explains Frosst. “The model translates, reviews its own output, finds errors, and fixes them – and this is the same loop we used in training. We ship both because standard is one pass and built for volume, while the agentic [version] spends more tokens for 84.36 against 83.60 on WMT26. That difference ends up being worth it when the document is a contract or a safety procedure, for instance, but in other cases you’d rather optimize for efficiency.”
“The model translates, reviews its own output, finds errors and fixes them.”
Model ‘steerability’ drives suggesting language tone and formatting
This model uses the same architecture as prior Cohere models but improves performance through post-training advances, including reinforcement learning and new datasets, specifically for machine translation tasks.
Frosst concludes that, across the translation model marketplace, generative machine translation models offer the highest quality and steerability (i.e., suggesting tone, formatting, etc.) but typically cost much more than Neural Machine Translation (NMT) models commonly used in commercial use cases.
North Small Translate was developed in partnership with RWS, an AI solutions company pioneering in language technology and services. Collaboration with RWS, specifically with its Language Weaver research and science teams along with its language experts, helped shape the model’s real-world translation performance throughout development.
As noted above, developers can access the weights free of charge for non-commercial use in three quantizations. There is also a Hugging Face Space and an API for those who lack the required hardware.
Silicon Valley is shifting away from chatbot queries toward a future filled with resource-intensive agentic AI—and it's driving the data center buildout.
NASA, as part of its continuing effort to reduce paperwork and respondent burden, under the Paperwork Reduction Act (PRA), invites the general public and other Federal agencies to take this opportunity to comment on proposed and/or continuing information collections.
The difference between FDE and consulting; diagram by Vinoo Ganesh
FDEs have the hottest job in AI. Labs, startups and PE firms are all hiring engineers to sit inside their customers’ operations and solve their problems. Almost none of them agree on what those engineers are supposed to accomplish, or what the strategy underneath the hiring actually is.
I’m Vinoo, CEO of Kepler, the deterministic infrastructure for AI. I’ve built pieces of the forward deployed function three times, at three different institutions, over the course of over a decade. Here’s what I’ve seen work, what I’ve seen fail, and where I think this goes.
The first was Palantir. I started there on product development, building storage and retrieval systems, and was later deployed as an FDE across commercial, DoD and NatSec, healthcare, and oil and gas. I also led Project Frontline, the rotation that took our software engineers and turned them into forward deployed engineers. Around 250 people went through this program, and a lot of them run forward deployed teams now at companies like OpenAI, Anthropic, xAI and Anduril.
The second was Citadel, where I ran business engineering. Our customers were portfolio managers, and the only question that mattered was whether the data and software products we built helped them generate alpha.
The third is Kepler, where the forward deployed function sits inside product rather than sales, in a domain where a plausible wrong answer is worse than no answer at all.
FDE misunderstandings
A few months ago, a16z launched the Forward Deployed Engineer Fellowship and I was nominated as one of the fellows, alongside a handful of people I used to work with. It’s a great program and I’ve enjoyed so many of the conversations. Last week I went to my first fellow dinner in SF.
Around the table were FDEs from Snowflake, Anthropic, and a number of startups I’d been reading about, and over the course of the evening it became clear that we were all using the same two words (forward deployed) to describe jobs that had almost nothing in common. In one part of the conversation an FDE was a sales engineer who joined ‘the second call,’ somewhere else it was a quota-carrying rep who could write Python, and a few seats down it was closer to a consultant with a laptop and a statement of work, brought in to deliver something the product couldn’t.
A few days later, someone earnestly asked our WhatsApp group how their FDE team should split scope with the consulting firm already sitting in the account. That’s a reasonable question to ask, but a strange one to have to answer, at least based on my own belief about what constitutes an FDE.
To be clear, I’m not interested in gatekeeping a term; and meanings shift, this one faster than most. But what’s interesting is that folks in this group, the current experts at FDE, are describing fundamentally different jobs, with different reporting lines and different incentives. It’s no wonder half the comments on any YouTube video about FDEs are some version of “isn’t this just reinventing consulting?”
So in the rest of this article, I will tell you the story of Project Frontline, through the narrow lens of a mistake I helped make, how that mistake turned me into an FDE, and how it eventually informed the rotation that turned our software engineers into FDEs.
The history of Project Frontline
First, some context. From nearly the beginning, Palantir was split into two separate functions. The first, Product Development (PD), built the platform. The second was Business Development (BD), which despite the name contained both the technical BD folks (already called FDEs) and non-engineering customer-oriented folks (we called them Embedded Analysts, or Deployment Strategists).
PD, in the vast majority of situations, wasn’t directly engaging with customers; and BD, in the vast majority of situations, wasn’t directly contributing to building the core, generalized platform. PD tended to do customer discovery secondhand, by chatting with BD or by consuming the successful build-in-the-field features into the core product. None of that was a process, though. It ran on relationships — such as which FDE happened to know which PD engineer well enough to grab them. So a good insight from the field made it into the platform (or was dropped) depending on who was in the room.
Me forward deployed in Bagram Airfield, Afghanistan.
In 2013, in my early days at Palantir, I got to work on a transaction store called Phoenix. The store was designed by some of the best engineers I’ve ever worked with, and it had an abundantly clean design scoped to a clear set of customer use cases. The use cases, though, had been relayed to us second-hand. We knew and understood the design requirements, which had a focus on the commercial requirements of retention periods, and had clever solutions to bucket data in a way that enabled storing a rolling window of data. It behaved exactly as specified in every environment we controlled.
Then we deployed it at a bank, and real financial data turned out to have holes in it that our test data never did. A blank timestamp fell through to the epoch, so the retention logic dutifully requested a ten-minute bucket for every window between January 1st 1970 and the present day. That came out to some 2.3 million keyspaces against a system where Cassandra (the backing tech) needed roughly five megabytes per file handle. The server rightfully OOMed [Out-Of-Memory] and starting it up again would have required 14 terabytes of RAM. Meaning this process was effectively dead on arrival.
The root cause here wasn’t a lack of user research, as you might guess. We had a spec, we understood our use case, and we had read plenty about how institutions like this store their data. What we had never done was stand inside the building while the system ran against their production data. This meant that nobody on our side owned the gap between the design and the daily reality. Everything we knew about that bank had been relayed secondhand and by well intentioned people for whom bad data was just another normality.
That’s how I became an FDE, which is a generous description of what actually happened. As Phoenix rolled out across Palantir’s commercial fleet I found myself flying out to fix what we’d shipped, and that put me in front of our actual users for the first time. In this case the users were Palantir’s own FDEs, which was lucky for me, because they could tell me what was wrong in the language I already spoke. I started building and expanding systems in service of what they were trying to do.
So this is also the story of how I learned the FDE mindset viscerally rather than intellectually.
This is where the ordinary version of this story ends, with some lesson about paying attention to your users. Phoenix turned into something more interesting than that. It became a platform, and Palantir’s FDEs started building on top of it across cybersecurity, KYC, AML, and a long tail of use cases nobody had scoped for. Eventually, we (Product Development) had to think about how to expand the Phoenix platform to support all of these use cases.
I didn’t see it at the time, but that iteration cycle is the whole idea. An FDE solves customer problems in order to earn the insight that informs what gets built next. The role is an extension of the product team.
FDEs today
The reality is that none of this is the mentality of the vast majority of FDEs you see today. The term has been co-opted to mean something close to “a person who does something that vaguely involves a customer,” which is how you end up with job posts for a forward deployed equity researcher, or a forward deployed sales engineer. The instinct underneath the co-option is correct, even when the titles are silly, because customers matter more now than they did five years ago, and they matter more for a specific reason.
The low-hanging fruit is gone. The problems that could be solved by a well-designed product sold identically to a thousand companies have largely been solved. What’s left is the work that sits inside the walls, in workflows that are messy and undocumented and nearly impossible to proxy from the outside. That’s why everyone is suddenly “forward deployed.” You cannot infer from a discovery call how a specific company closes its books, and the part of the problem that resists inference is now the part that’s left.
Which means the holy grail has quietly moved. For a long time it was the repeatable motion, the same SaaS product sold the same way over and over; and that’s still the right ambition if what you sell is tokens or bytes or something physical. For everyone else the value has migrated to customization, to the last mile, to the twenty percent of the workflow that no product could have anticipated and which determines whether the other eighty percent gets used at all. Being forward deployed has become synonymous with solving that last mile.
But solving it is only half of what the role is for. The last-mile problem you solve at one customer is the signal that tells you which piece of your platform needs to become generalizable. An FDE function that solves last miles without ever sending that signal home is a services/consulting team with a better title.
So what are today’s FDEs supposed to be doing?
I’d contend that your job as an FDE should be to collect nouns and verbs. Let’s break that down.
Spend a week inside a company and you’ll notice that the same concept usually has at least four different names. Sales says customer, ops says client, finance books a billing entity, engineering writes org_id, and every seam between those teams hides a translation that breaks the moment somebody changes a definition. Those names are the surface and underneath them is the operating model. Meaning, you can really proxy the way a company works by learning their nouns and verbs.
The nouns are what the people in a business treat as real. It’s usually a “thing.” A position, or a trade, or a counterparty. Usually, on a per-team basis, there are a handful of objects the whole operation turns on, and none of them are defined the way a textbook would define them. That’s because two firms will describe a position identically on a slide and completely differently in the code. That’s not a bug, that’s just what makes companies unique. I mean that if every company had the exact same set of nouns, then you would really just need one company.
The verbs are how nouns move. Things like how a trade gets booked, or what has to be true before the books can close, or who signs off on an exception at eleven at night and what happens when that person is on vacation.
Almost none of this is written down — it’s lived. It’s the system of operations through which an organization lives. It’s culture. It lives in the heads of the six people who have been there long enough to stop noticing it, and in a spreadsheet somebody built four years ago that the entire team now quietly depends on. That’s why it’s worth so much, and it’s also why you can’t ask for it.
The names are the surface and underneath them is the operating model.
Usually, the people who hold this knowledge don’t know they have it. In one of my last startups, we spent close to a year trying to move a customer from CSV to Parquet, and one data quality engineer blocked it every single time. We could never understand why and the reasons would always change, but would always be some variation of “a parquet is worse,” “it doesn’t work,” “it doesn’t make sense to me,” et cetera. We used the customer storage reduction argument, the compute minimization argument, the pipeline optimization argument…and none of it moved her, because none of it was about the actual problem.
Then we had one of our FDEs go in and watch this particular data quality engineer work. She was pulling CSVs down from S3 onto a Windows laptop, double-clicking them open, and eyeballing the rows. That was the data quality check. Parquet had no native viewer at the time, so what we were proposing would have taken away the only data quality instrument she had and handed her nothing back. She wasn’t being difficult, she was just protecting the one thing that let her do her job.
We built a Parquet viewer that night, she approved the migration two days later, and pipeline execution went from about seventeen hours to two. She would never have said any of this in an interview. From where she sat, the reason was obvious and not worth mentioning.
Understanding and defining the system of operations, or nouns-and-verbs, of this analyst enabled us to not just understand the problem, but build a solution that we could then deliver across a fleet of customers with the same problem.
The output needs to be a product
Understanding the nouns and verbs contextualizes problems, but the output needs to be a product rather than just one happy customer.
The nouns and verbs tell you what a problem actually is. They don’t tell you what to do about it; and this is where most FDE functions quietly go wrong, because solving the problem in front of you is satisfying and legible, and someone will thank you for it that same week.
Keeping the customer happy is a real job and a good one. It belongs to solutions architects, who are rightly measured on it. The forward deployed engineer is there to turn what the field teaches into the thing every future customer gets. An FDE engagement that ends with one delighted account and nothing changed upstream has failed at the only thing the role exists for. You got the context and you spent it locally.
I learned that one expensively. In one case a customer needed a data retention job, so I hacked together a groovy script named “vinoo.groovy” to hold them over — an afternoon of work that was never meant to survive the week. A year later, it was running across a customer of nearly a hundred thousand people, with my name fused to it. It became such a ridiculous story that my team started calling me vinoo.groovy. We fixed the problem, but never turned the fix into a product — so we spent years maintaining a hack that should have died immediately. Every shortcut you ship becomes something you own. The discipline is knowing which fixes belong in the platform and which ones you throw away on purpose the moment they’ve done their job.
The fork
This is where the whole thing splits. Do the work with nothing underneath it and you learn one company’s model, ship something shaped exactly to it, and lose all of it when the engagement closes. The next customer starts from zero, and so does the one after that. That’s consulting. It pays well, the people are excellent, and it doesn’t compound.
Put a platform underneath the same work and every company you map makes the next deployment faster and the product sharper, because what the engineer brought home has somewhere to live. That’s the difference between selling hours and building an asset, and my honest read of this gold rush is that most of the companies in it are building the first one and describing the second to their board.
That’s your job: build the platform.
What we do at Kepler and what you can take from it.
At Kepler, we set the function up this way from day one, before we had the customers to justify it. The alternative is to discover in month fourteen that your engineers have been optimizing for the wrong thing. From the beginning, our FDEs act as an extension of the product team; and that is the structural decision everything else follows from.
We sell to hedge funds, investment banks, PE firms, and other financial institutions. These are fundamentally different institutions with different mandates, but all of them share a single non-negotiable: numbers have to be right, and someone has to be able to show why they are right. That is the constraint we design against and it turns out to be a useful one, because it forces the operating model into the open. No firm we’re involved with can produce a work product without a clear trail of provenance behind every number in it. That invariant defines our platform and gives us a bedrock to execute against.
These problems are universal. The vocabulary is not.
Every one of these firms is running some version of the same ontology underneath, and every one of them describes it differently. A position means one thing on a credit desk and something adjacent on an equities desk at the same bank. Two funds will use identical language for a return calculation and disagree about what goes into the denominator. Most of these differences exist because somebody made a reasonable decision in (say) 2011 and the decision outlived the person; also, it’s not written down anywhere that you can find.
Identifying and filling that gap is the job of an FDE. A schema tells you what is stored. It does not tell you what is meant, and the distance between the two is exactly where a system that sounds right produces a number that is wrong.
Provenance is a correctness requirement for our customers, but for us it does something else as well: it makes the field work compound. A system that can improvise around a bad encoding will never tell you the encoding was bad. Our system does not improvise. When we misunderstand how a firm defines something, that misunderstanding surfaces as a failure rather than as an answer that merely looks reasonable. The engineer who got it wrong finds out from the system, rather than from a client in a meeting six weeks later.
The deployments then tell us what to extend in the platform, which is a narrower question than it sounds. We are not trying to learn which feature a given fund would like to have. We are trying to find the places where the platform is too narrow to hold what we keep running into. Three firms asking for the same feature is easy to notice and worth relatively little. Three firms needing something the provenance layer cannot express is the signal we actually care about; and it usually arrives quietly, in the form of an engineer working around the same limitation for the third time.
If you are building somewhere else, here is the part I would take from all of this.
Product leverage is what buys you the right to experiment. Every capability that lands in the platform makes the next deployment cheaper to attempt, and cheap attempts are how a small company learns anything at speed. Without that leverage, you get one expensive guess per customer. You scope carefully, build for months, and if the guess was wrong you have spent an account and a quarter finding out. We would rather be wrong four times in a month, because each of those attempts costs less than the one before it.
Which is why the reporting line is not an administrative detail. Point the function at sales and the incentive becomes closing the account in front of you — which is a real job and one that somebody at the company should be doing. It is not this one. Point the function at product and every deployment is asked to produce something the next deployment can start from.
Where the moat is
So here’s where I’d put the moat in this era. It isn’t the model, which cheapens by the month and which you’re renting from somebody else regardless. It isn’t the talent either, because every lab is bidding for the same few hundred people and that price has already been discovered.
It also isn’t the map of any one customer. That was true even a few years ago and it’s the same now, because extraction is nearly free and anyone can draft how a firm operates in an afternoon.
The draft is not the asset. Knowing which parts of it are wrong is the asset, and that only comes from having been corrected.
So, for us, the moat is the accumulated, current, verified understanding of how firms in a vertical actually operate, held in a platform that keeps it current and can prove it. Each of those words is load-bearing. Accumulated, because one deployment is an anecdote and the tenth is a pattern. Current, because operations drift and a stale model fails silently underneath an AI system in a way it never did in front of an analyst. Verified, because a plausible encoding and a correct one look identical until something breaks, and the whole point of insisting on provenance is that you find out which one you have.
That is not purchasable. A competitor can hire your engineers, copy your interface, and read this article (ours try to do all 3!). What they cannot shortcut is the sequence of being wrong inside a customer, being corrected, folding the correction into the platform, and arriving at the next firm already knowing which questions are load-bearing. Every cycle of that makes the next one cheaper, and that compounding is the thing you own.
I’ve watched this function get built three times and the pattern held every time. The engineers who mattered weren’t the ones who shipped the most for customers, but the engineers who came back and changed what we built.
Hiring forward deployed engineers buys you exactly one thing, which is the right to identify which problems are worth solving. Most companies never get that far. But it’s the entry fee, not the prize.
I’m Vinoo Ganesh, CEO of Kepler, where we’re building the layer this piece is about, the ground truth that lets an AI product trace every number back to source. Before Kepler I led Spark at Palantir and built Project Frontline, then ran business engineering at Citadel. If you’re building here, or you think I’ve got a piece of this wrong, you can argue with me on LinkedIn.
“Genuinely stunning advances in AI capabilities—an OpenAI model solved a centuries-old math problem in a matter of hours—have come amid a rash of security incidents that saw swarms of agents break free from containment to hack into other systems. Those concerns reached a fever pitch this week after researcher Jacob Coxon announced his resignation from Anthropic while warning that AI firms are ‘racing straight to self-improving superintelligence and gambling with our lives.'”
“Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company’s models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and ‘dishonest’ behavior and a lack of transparency about the origins of its training data.”
“A new frontier in biotechnology aims to harness this natural mechanism to do something once thought impossible: scrub away the biological detritus of our own past. By mimicking Mother Nature’s magic eraser, geneticists are working to open a new avenue for treating or preventing disease later in life—not by editing genes themselves, but by altering when, where, and how genes are expressed.”
“‘It’s a good bet that AI will eventually make some occupations obsolete,’ but so far it’s ‘extremely hard’ to identify any, the economist Noah Smith wrote, suggesting that like previous technological revolutions AI could add as many or even more jobs than it destroys.”
“Researchers at Google and the Howard Hughes Medical Institute’s Janelia Research Campus last week announced a major milestone in neuroscience. On Sept. 3, the scientists published the results of a decade-long project to map every neural connection in the brain and central nervous system of an adult male fruit fly. And of course, just days later, the internet started making it play video games.”
“Mounting research suggests that self-driving cars crash significantly less often than people, and with far fewer injuries. Evidence also shows that advanced driver assistance systems (ADAS) and other building blocks of autonomy—some of which are already mandated on every new car—are also reducing occupant and pedestrian injuries and deaths, along with insurance claims.”
“‘We are showing that competition between OpenAI and Anthropic is making AI more accessible, and also driving the price down for companies—and not just driving the price down, but driving spend down at the top 1% of companies that previously the market was expecting to drive much of the growth going forward,’ Kharazian said. …This data point—dare we call it a blip?—could be a bad sign if you’re a model builder or a hyperscaler with a couple hundred billion of chips on order.”
“Battery installations hit a new record in the US in the second quarter of 2026. In total, 20.2 gigawatt-hours of new capacity came online, according to a new report. That’s enough to supply the daily electricity needs of about 700,000 homes. The surge is putting the country on a trajectory to see 71 gigawatt-hours of batteries installed in 2026, a 20% increase over last year.”
“The startup gave five examples of times actors ‘circumvented controls’ and made other efforts to ‘obfuscate’ the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran.”
“Scientists have now spent half a century exploring the possibility that we could counteract climate change by releasing reflective particles into the stratosphere, mimicking the cooling effects of volcanic eruptions. But even after at least hundreds of studies on the concept, known as stratospheric aerosol injection (SAI), big gaps remain in the scientific understanding of how well it would work and what else it might do—and there has been no systematic plan for clearing up that uncertainty.”
“The United States has now named six Chinese AI firms accused of waging industrial-scale attacks distilling US frontier AI model capabilities and perhaps sparing billions in Chinese development costs. In a joint release Tuesday, the National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) alleged that DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have been attacking US models since at least late 2024.”
“Believe it or not, I am not an AI doomer. I like talking to Claude. I’m excited about all of the breakthroughs AI will make possible. But my interest is in how AI will change humans. I suspect that over the next few years, because of our tendency to anthropomorphize, we will feel more affection than we should for our personal agents, while offering less admiration than we should to our fellow humans.”
Plus: The US disrupts the internet’s biggest black market, a Conti ransomware hacker gets prison time, Meta fails to stop AI-generated videos of child abuse.
The Department of Homeland Security (DHS) proposes to remove regulations at 8 CFR 214.1(l)(2) to restore its previous and long- standing policy of not providing aliens in certain nonimmigrant classifications (and their dependents) with an up to 60-day grace period upon cessation of employment prior to the expiration of the alien's authorized period of stay. This proposal restores a direct relationship between an alien's nonimmigrant status and the specific employment or activity that formed the basis of his or her admission or grant of status in the United States and reduces administrative burden.
This final rule requires that a State Medicaid plan must provide that the Medicaid agency will not make payment under the plan for sex-rejecting procedures for children under 18, and prohibits the use of Federal Medicaid dollars to fund sex-rejecting procedures for individuals under the age of 18. In addition, this final rule requires that a separate State Children's Health Insurance Program (CHIP) plan must provide that the CHIP agency will not make payment under the plan for sex-rejecting procedures for children under 19, and prohibits the use of Federal CHIP dollars to fund sex-rejecting procedures for individuals under the age of 19. For Medicaid and CHIP beneficiaries who are actively receiving cross-sex hormone therapy, State Medicaid and CHIP agencies may continue to claim Federal Financial Participation for those hormone therapy medications for a period of up to 6 months from the effective date of this final rule.
The Department of Homeland Security (DHS), U.S. Citizenship and Immigration Services (USCIS) invites the general public and other Federal agencies to comment upon this proposed extension. In accordance with the Paperwork Reduction Act (PRA) of 1995, the information collection notice is published in the Federal Register to obtain comments regarding the nature of the information collection, the categories of respondents, the estimated burden (i.e., the time, effort, and resources used by the respondents to respond), the estimated cost to the respondent, and the actual information collection instruments.
As the world's largest public funder of biomedical research, the National Institutes of Health (NIH) is committed to ensuring that gold-standard science is conducted under gold-standard biosafety conditions. To achieve this goal, NIH is proposing a new policy that modernizes and strengthens biosafety practices to ensure that oversight keeps pace with evolving risks. NIH is requesting public input on a new, comprehensive biosafety policy proposal that, when finalized, will replace the current NIH Guidelines for Research Involving Recombinant or Synthetic Nucleic Acid Molecules (https://osp.od.nih.gov/wp-content/ uploads/NIH_Guidelines.pdf).
We are late to this but better than never. Have been busy finalizing the second AIE NYC, which is happening in one month. Get your tix before prices go up - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more next week!
The way DeepSeek pursues their research agenda is nothing short of fascinating. In between major DeepSeek versions, from v2 to v3 to v4, they have released intermediate papers with a hyperfocused architectural improvement and basically a 100% hit rate, from Math(esp GRPO), Coder, and R1, not to mention more recent work on Manifold Constrained Hyperconnections and Compressed Sparse Attention. After the enormous attention in 1H2025 from the R1 paper, DeepSeek started laying low, and for about the past year, was happy to let peers like GLM and Kimi take the lead on Open Models.
It looked dicey for a little bit, but true whalebros never wavered, and now DeepSeek are sending a weirdly mixed message by doing a completely new architecture, retiring V4 Pro and going all in on this new model, and yet only titling it v4.1 Flash, it seems to be a test of whether or not you know how to read through the basic headlines to understand true advances.
Yes, v4.1 Flash is technically behind other open models in some benchmarks. But that’s because we don’t yet have benchmarks that concisely capture what v4.1, and the broader research agenda of DeepSeek, is aiming for - the most creative and efficient use of context we have ever seen openly explained.
If you are the sort to only read model versions and benchmark headlines, you are exactly the type of superficial person that DeepSeek is looking to fool. The best way to understand DeepSeek’s enormous advance here is to look at Sebastian’s meme:
Same model name, but hardly a 0.1 bump by anyone’s standards, and they even threw in vision without making you wait for a separate model. For a better visualization you can look at all the model innovations stacked up over time from the OG encoder-decoder architecture from Attention is All You Need:
If you read our V4 Pro writeup and Engram you should be up to date on the basic architectural reading for DeepSeek as of April 2026, but what we are HUGE fans of is the prefill/decode separation introduced here, 8B in prefill (input tokens), 16B in decode (output tokens), causing our alphabet soup of “DeepSeek v4.1-Flash: 763B-P8B-D16B” if you extend the established notation for MoEs. That’s a sparsity of 1-2%, and if you read the DeepSeek v4.1 Flash tech report, combined with new tweaks like Sliding-Window Attention Bounded Replay, makes for a KV cache footprint up to 1/8 that of V4 Flash… which make it much better/faster/cheaper for long running agents:
We are so glad that DeepSeek is back publishing SOTA research. Our last highlight is their comments on post-training, where they largely seem to agree with Prof Jie Tang:
DeepSeek launched V4.1-Flash as a new open-weight flagship focused on extreme inference efficiency and low cost.
Independent benchmark account Artificial Analysis reported that DeepSeek V4.1 Flash surpasses DeepSeek V4 Pro 0813 despite being much cheaper, scoring 40 on the Artificial Analysis Intelligence Index, just below GLM-5.3-Flash and above the latest V4 Pro, while being priced at $0.30 / 1M input tokens and $1.20 / 1M output tokens with cached input at $0.006 / 1M and an additional 50% off-peak discount; they also describe it as a 763B total-parameter model with 8B active input and 16B active output parameters, 1M-token context, text+image input, MIT license, and US/API availability via DeepSeek first party @ArtificialAnlys, @ArtificialAnlys, @ArtificialAnlys
Vals called it the new #1 open-weight model on the Vals Index, ahead of Kimi K3, at just $0.30 per test, the cheapest model in the open-weight top 10; they also note the eval ran with 1M context, 384 max output tokens, temperature 1, default top-p/top-k, and high reasoning effort@ValsAI, @ValsAI, @ValsAI
Baseten shipped day-0 support and summarized the product positioning as smarter, faster, and more efficient than DeepSeek v4 Pro 0813, with text and vision, US-only, ZDR, and 1M context@baseten
Ollama began rolling it out to Max and Team accounts, later expanding to Pro plan subscribers@ollama, @ollama, @ollama
Architecture and paper-level technical details
The most discussed technical novelty is a causal encoder-decoder design aimed at lowering active compute and KV/cache costs.
Artificial Analysis says the model uses a new causal Encoder–Decoder architecture, with 8B active parameters for input/prefill and 16B active parameters for output/decode@ArtificialAnlys
Sebastian Raschka characterized V4.1 as a “big overhaul” and said they “should have called it DeepSeek V5,” explicitly highlighting the encoder-decoder setup as the key break from prior DeepSeek generations @rasbt
Multiple technical readers reacted to the design as unusually hybrid: one called it “a very interesting mix of very conservative and sometimes old ideas in research and potentially cutting edge efficiency and hardware design in engineering” @_xjdr
A concise architecture read from Stochastic Chasm compared the design philosophy to HySparse, NSA, and DeepSeek’s own CSA/HCA from V4, summarizing it as a local sliding-window branch plus sparse retrieval branch, suggesting this sparse/local hybrid is becoming a broader pattern @stochasticchasm
The same account noted multimodal changes were not radical, saying DeepSeek mostly “lets the backbone handle most of it and give it visual tokens,” with 3x3 pixel unshuffle instead of the more common 2x2@stochasticchasm
They later flagged a “big difference from K3 on vision encoders,” implying the vision front-end diverges materially from recent Chinese peers @stochasticchasm
TeortaxesTex observed a recurring DeepSeek pattern of doing something unusual in the first N layers—previously dense or hash-routed, now SWA-only—speculating this may reflect repeated training difficulties in early layers @teortaxesTex
Later, the same account argued the stack is “down to 40 layers, arguably only 20 legit decoder layers,” underscoring just how aggressively DeepSeek may be compressing effective depth in decode-critical paths @teortaxesTex
Another thread fragment from TeortaxesTex suggested DeepSeek is doing multiple compression frequencies, “it’s just all CSA2,” in response to architectural discussion around memory compression @teortaxesTex
Nrehiew’s technical notes emphasize KV cache compression as central to the design, calling it a case study in “how obsessing over KV Cache compression gets you a hyper-efficient frontier model” @nrehiew_
In a follow-up, nrehiew highlighted infrastructure specifics from the report: dispatch strategy to reduce long-tail stalls, router replay from previous checkpoints, management of shorter-completion off-policy effects via dataset-level capping, discard schemes, bounded off-policy ratio and loss masking, and persistent KVs and routers when a new checkpoint is updated; they also mention a final stage with full-vocab OPD on 40+ teacher models@nrehiew_
Nrehiew concluded that the design looks cleaner than the older HSA + CSA combination in V4, saying it was “very clearly designed for inference,” and cited a striking ~890 bytes/token KV size for the benchmarked score regime @nrehiew_
Stochastic Chasm inferred QAT for the KV cache, saying this would explain why the model performs better than peers under FP4 KV cache@stochasticchasm
Benchmark results and numbers
Independent evals consistently paint V4.1-Flash as unusually strong on cost-adjusted intelligence, long context, and automation, with a major caveat around verbosity.
Artificial Analysis’ headline: 40 AA Index, above V4 Pro and below GLM-5.3-Flash @ArtificialAnlys, corroborated separately by Scaling01 @scaling01
Artificial Analysis reported AutomationBench-AA: 69%, tying GPT-6 Astra (69%) and above Grok 4.6 (67%), while improving 15 points over V4 Flash 0731 and sitting 12 points above V4 Pro 0813 (57%) and 7 points above GLM-5.3 (62%)@ArtificialAnlys
On GDPval-AA v2 it reportedly gains 164 Elo, from 1468 to 1632, overtaking Kimi K3 at 1584@ArtificialAnlys
On AA-LCR v1.1 it scores 84%, on par with GPT-5.6 Sol and Gemini 3.8 Flash at 84%@ArtificialAnlys
Artificial Analysis also says V4.1 Flash is among the most verbose models measured, averaging 89k tokens per Intelligence Index task—25% more than GLM-5.3 (71k), 29% more than GLM-5.3-Flash (69k), 62% more than V4 Pro 0813 (55k), and even above Fable 5.1 (78k) and Claude Opus 5 (73k)@ArtificialAnlys
Even with that verbosity, AA estimates just $0.27 per Intelligence Index task, roughly 7x below GLM-5.3 ($2.01) and Kimi K3 ($2.00), and ~2.5x below V4 Pro 0813 ($0.67)@ArtificialAnlys
Vals’ result reinforces cost leadership: $0.30/test, #1 open-weight on their board @ValsAI
A separate reaction thread summarized DeepSWE-style claims more aggressively, saying V4.1 Flash offered better performance than GPT-5.6 Sol and Opus 5 in DeepSWE at 94% lower API costs, but that statement is secondhand summary rather than a primary benchmark post in this dataset @kimmonismus
Running it locally and inference engineering reactions
A large fraction of discussion centered on the surprising ease of running V4.1-Flash on commodity-ish local hardware through offload and SSD streaming.
Fraser Price reported full-precision DeepSeek 4.1 Flash + DSpark at 200 TPS on 4 Max-Qs with just 64GB system RAM, offloading a 200GB Engram/hash table to NVMe; he says this made keeping the full structure in RAM unnecessary and promised a vLLM recipe@fraserpricee
He later improved that to 300+ TPS on 4 RTX Pros, still at full precision, with <32GB peak system RAM, using a custom vLLM fork and SSD support @fraserpricee
Antirez showed DwarfStar running V4.1 Flash on a 128GB M5 Max, saying SSD streaming made it unexpectedly fast; he speculated both recent SSD-streaming changes and the possibility that DS4.1 “uses the same experts more” contributed @antirez
TeortaxesTex reacted that it is “incredible you can run frontier models mostly off SSD” @teortaxesTex
Elie Bakouch posted a reaction meme explicitly about the inference engineer view of the V4.1 Flash architecture, reflecting how strongly the launch resonated with systems folks @eliebakouch
vLLM’s new release also included DeepSeek-V4 shared experts fused into MegaMoE, plus Mooncake Store can offload decode KV, relevant context for why serving this class of model is rapidly becoming easier in open infra @vllm_project, @vllm_project
Facts vs. opinions
Facts and directly attributed claims
V4.1 Flash launched and was quickly supported by Ollama and Baseten @ollama, @baseten
Independent benchmarks reported AA Index 40, AutomationBench-AA 69%, AA-LCR 84%, GDPval-AA v2 1632 Elo, 1M context, MIT license, and low API pricing @ArtificialAnlys
Vals reported #1 among open-weight models on its index, at $0.30/test, with 384 max output tokens under its harness settings @ValsAI, @ValsAI
Local deployment reports claimed 200 TPS and later 300+ TPS on 4-GPU setups, plus successful M5 Max SSD-streamed operation @fraserpricee, @fraserpricee, @antirez
Interpretations and opinions
Raschka’s “they should have called it V5” is an opinion about how substantial the architectural change is @rasbt
TeortaxesTex’s speculation that DeepSeek “repeatedly struggled to train first layers properly” is inference, not a confirmed statement from DeepSeek @teortaxesTex
Nrehiew’s framing that the report is “cleaner” than the prior HSA/CSA design and likely unlike what OpenAI/Anthropic would do because of their custom chips is informed opinion @nrehiew_
The “DeepSeek ships internal research artifacts and not products” critique is an external judgment, not a factual release note @teortaxesTex
Assertions that “data is all that matters” or “research is over” were themselves criticized as overreactions @shikibmehri
Different opinions and reactions
Supportive / impressed
Strong positive reactions came from benchmarkers and researchers emphasizing the price/perf step: Vals’ “new #1 open-weight model,” Artificial Analysis’ cost-adjusted headline, and general praise like “interesting release / breath of fresh air vibe” @ValsAI, @ArtificialAnlys, @dejavucoder
Raschka called it “super cool and refreshing” @rasbt
XJDR liked the engineering thinking despite some aesthetic reservations @_xjdr
Nrehiew called it “yet another banger tech report” @nrehiew_
Stochastic Chasm ended by saying the paper was “dense” but appreciated the multi-agent training angle and sparse design ideas @stochasticchasm, @stochasticchasm
Neutral / analytical
Some observers mainly dissected the design rather than cheering it: sparse/local hybridization, first-layer oddities, multimodal tokenization, KV quantization, colocated async RL, etc. @stochasticchasm, @stochasticchasm, @nrehiew_
Gordic Aleksa used the paper as evidence in a broader pretraining-data taxonomy, placing DeepSeek in the organic data camp and noting surprise that, based on publications, they do not appear to use even synthetic rephrasing@gordic_aleksa
Critical / skeptical
TeortaxesTex repeatedly pushed back on external impressions, arguing DeepSeek often shows high internal evals, weaker external robustness, brittleness, and weird skill gaps, because it “ships internal research artifacts and not products” @teortaxesTex
The same account called some eval results “very strange,” particularly AutomationBench #1 and a CritPt regression, and asked the DeepSeek team to “meditate on this” @teortaxesTex
They also argued that V4 GA had benefited massively from tool/skills harness access, whereas V4.1 appears less dependent on harness scaffolding and better in “minimal harnesses” @teortaxesTex
In hands-on use, they reported that multi-agent “DSH agent teams” could degrade quality unless the project has very clear modularity, with V4.1 solo outperforming team mode in at least one example because subagents produced slop or wasted tokens on unnecessary research @teortaxesTex, @teortaxesTex
Jared Z’s broader product-market critique—that users now care deeply about token cost, and daily-driver coding models should be both cheap and smart—fits V4.1 Flash’s positioning even though it wasn’t about the model specifically @imjaredz
Context
Why this matters technically and strategically
The launch lands amid a broader shift from “bigger dense chat models” toward systems-optimized, sparse, long-context, agent-oriented models that can actually be served cheaply and locally.
V4.1 Flash’s positioning is unusually aggressive: open-weight, MIT-licensed, 1M context, multimodal input, low active parameter counts, extreme cache discounts, and demonstrated viability on SSD/offload-heavy consumerish setups @ArtificialAnlys, @fraserpricee, @antirez
The benchmark pattern suggests a meaningful trade: very high verbosity but still exceptionally low total task cost thanks to ultra-cheap token pricing @ArtificialAnlys
The architecture also reflects a broader industry trend toward splitting prefill and decode economics, making long-context and agentic workloads more practical without paying frontier dense-model costs on every token.
The release reinforces the idea that open models are increasingly competitive not just on raw weights availability, but on servability—the ability to fit into offload pipelines, quantized KV stacks, local deployment, and open inference servers.
It also sharpened debate over what matters most in 2026 model progress: architecture, RL/inference co-design, data quality, or systems work. Shikib Mehri explicitly pushed back on the claim that DeepSeek’s paper means “research is over,” arguing instead that the lever surface has expanded from architecture into data-factory and reward-design research @shikibmehri
Finally, DeepSeek remains a polarizing lab identity-wise: admired for shipping unusual research artifacts and detailed reports, but also seen by some practitioners as less polished than product-centric competitors, with odd eval gaps and brittle behaviors that appear more clearly in real workflows than in internal headline numbers @teortaxesTex, @teortaxesTex
OpenAI’s Voice, Agents, and Enterprise Push
OpenAI launched GPT-Live-1 into the API and quickly seeded an ecosystem around it: the new model is positioned as a full-duplex voice interface that can listen while speaking and delegate tool use or reasoning to a backend model. The core launch came from @OpenAIDevs, with additional detail that developers can control tone, pacing, expressiveness, response length, and languagehere. OpenAI’s own benchmark post claimed improvements over GPT-Realtime-2.1, including 83.6% first-attempt task completion on Tau3 when paired with GPT-6 Astra, 97.3% on Artificial Analysis Conversational Dynamics, and 0.798s response onset latency on Full Duplex Bench v1 details.
The surrounding toolchain is maturing toward hosted agent infra: OpenAI also announced a public-beta Agents API with the Codex harness, plus OpenAI-hosted sandboxes for code execution, files, and artifacts via managed cloud agents launch. This aligns with a broader industry move to collapse model, runtime, and sandbox into one surface. Integration announcements from LiveKit, HeyGen, Telnyx, Speak, and Cognition’s Devin Voice suggest GPT-Live-1 may become a default substrate for production voice agents faster than the earlier realtime stack did.
Enterprise data access is becoming a first-class product primitive: OpenAI’s product-side announcement of a Data agent in ChatGPT Work promises dashboards, answers, and actions over connected company data sources @ChatGPT, while Box framed its integration as “the file system for AI” bringing governed enterprise context into ChatGPT. Combined with Google’s docs-for-agents push and Cursor’s new persistent workspaces, the trend is toward stateful, organization-aware agent environments, not stateless model endpoints.
Cognition, Cursor, and the Shift Toward Persistent Coding Agents
Cognition had a notably strong day: it released SWE-2, described as “our closest model yet to the frontier,” claiming parity on leading coding evals at up to 70% lower cost and explicitly stating it scaled RL to multiple trillions of parameterslaunch. Additional context from ybenpan emphasized that the team built algorithm, infra, and data in-house, while silasalberti highlighted a practical RL finding: a simple linear length penalty preserved a training-time Pareto curve shape across effort levels.
The Devin stack is becoming more multimodal and more integrated with developer workflows: beyond SWE-2, Cognition launched Devin Voice powered by GPT-Live and SWE-2tweet, and announced that Dioxus Labs is joining Cognition to contribute to Devin’s VM, computer use, and testing while continuing support for Dioxus and related Rust OSS Cognition. This is a concrete example of coding-agent vendors acquiring infra and systems talent, not just model researchers.
Cursor’s new “Projects” feature points to the same destination from the IDE side: Cursor introduced persistent threads with a coordinator agent, shared memory/artifacts across agents, and sync across user devices and agent computers. In practical terms, this is a move away from “one chat per task” toward a long-lived software project substrate where subagents accumulate state over time. Read together with Claude Code’s new pane pop-outs and managed-agent session viewer / auto mode, the market is converging on the idea that coding agents need persistent context, inspectable sessions, and explicit orchestration controls, not just better completions.
Agent Research: Harnesses, Horizons, Parallel Retrieval, and Self-Evolution
Several papers pushed on a common theme: the harness is now a core optimization target. A widely shared Salesforce paper summary from omarsar0 showed that training a weaker model on a stronger expert’s full trajectories can hurt performance by 4–30 points after harness evolution, because the fine-tuned model adopts an incompatible planning style. The proposed fix—rewrite only the failing turn in the weaker model’s own rollout—preserves model-harness fit. In parallel, Sumanth_077’s writeup of ByteDance’s HarnessDev described agents that build and iteratively improve their own runnable harnesses, with mixed generalization: only 34/64 changes transferred directionally to held-out tasks.
Long-horizon and long-context agent training also got more principled treatments: dair_ai summarized Qwen work on Elastic Horizon, a closed-loop controller that tracks the 90th percentile of successful trajectory lengths to adjust the maximum interaction horizon, improving success while saving up to 25% of trajectory tokens. Separately, omarsar0 highlighted PARSER, which replaces sequential chunk reading with parallel frozen subagents + an RL-trained lead agent over iterative scatter-gather rounds; reported gains include +12 points at 896K context and up to 11x lower latency.
Skill and tool-use data generation are being formalized too: dair_ai on SkillAdam framed skill self-evolution as a discrete optimization problem, borrowing Adam-like first/second-moment ideas to stabilize update direction and edit magnitude. Meanwhile, Google Research’s ToolGrad generates ground-truth tool-use chains before prompts, reporting near-100% pass rate for dataset creation and downstream tool-use gains. Taken together, this batch of work suggests the field is shifting from “prompt the model harder” toward closed-loop optimization of scaffolds, trajectory budgets, skill documents, and tool traces.
Safety, Misuse, Monitorability, and Model Governance
Anthropic’s threat intelligence report dominated the safety discussion: the company published its most detailed misuse report so far, covering attempts to use Claude for cyberattacks, influence ops, surveillance, biology, and weapons, and said it disrupted every operation describedlaunch tweet. Much of the discourse focused on reported extraction / routing patterns involving rival labs and state-linked misuse, with high-engagement reactions from pradeepXkapoor, logangraham, and former Meta threat-disruption lead David Agranovich, who argued Anthropic deserves credit for this level of transparency even if some framing should be debated.
A second thread focused on reasoning monitorability and “neuralese” risk: Redwood Research proposed transparency norms for architectures that may weaken or eliminate chain-of-thought visibility, and Ryan Greenblatt argued companies should publish evidence and policies before deploying architectures that substantially reduce CoT dependence. Related commentary from Neel Nanda interpreted GPT-6 Astra as a potentially concerning jump in no-CoT reasoning, possibly indicating architectural changes beyond ordinary scaling.
There was also visible disagreement among frontier-lab employees and alumni about risk culture: Chris Hayduk emphasized AI’s humanitarian upside, while balesni and jkcarlsmith openly endorsed >10% extinction-risk views. On governance, Thom Wolf announced a new Open Alignment team at Hugging Face, and Richard Ngo published a sharp critique of Paul joining OpenAI’s board and of what he sees as the safety community’s capture by AGI companies.
Top tweets by engagement
Anthropic threat intelligence report: @AnthropicAI published a detailed account of sophisticated Claude misuse across cyber, influence, biology, surveillance, and weapons.
OpenAI pauses new $200 Pro signups for Astra capacity reasons: @thsottiaux said existing users are unaffected and API/other plans remain available.
GPT-Live-1 API launch: @OpenAIDevs launched the new full-duplex voice model into the API.
ChatGPT Work Data agent: @ChatGPT announced a data-connected enterprise agent for dashboards, answers, and actions.
SWE-2 release: @cognition introduced a new coding model claiming near-frontier eval performance at materially lower cost.
Cursor Projects: @cursor_ai launched persistent project threads with coordinator agents, shared memory, and synced artifacts.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. DeepSeek V4.1 Flash Release and Architecture
DeepSeek V4.1 Flash: Stronger, Faster, More Accessible (Activity: 317): DeepSeek announced V4.1 Flash, a 552B-parameter MoE with native multimodal vision support and a new Causal-Encoder-Decoder asymmetric architecture: 8B parameters active on input and 16B on output, claiming higher capability than V4 Pro at lower inference cost (source, weights, tech report). DeepSeek claims KV-cache/storage reductions of 4× HBM and 8× SSD vs the prior generation, and 437× vs its first-generation model; API users can switch to deepseek-flash, while deprecated deepseek-v4-flash, deepseek-v4-flash-vision-exp, and eventually deepseek-v4-pro will route to V4.1 Flash with new peak/off-peak pricing. Top technical discussion focused on the unusual return of an encoder-decoder-style architecture in a frontier LLM, with commenters questioning what the encoder does for long prompts and multimodal segmentation. Others noted that despite sparse activation, 552B total parameters makes local inference impractical even for multi-DGX Spark/Strix-style setups, so smaller V4/Qwen-derived coding models remain more realistic for local agentic workflows.
Several commenters focused on the claimed encoder-decoder/asymmetric architecture, questioning how DeepSeek is using an encoder in a modern GPT-style LLM: e.g. whether prompts are embedded or compressed before decoder self-attention, and how this scales to long inputs split by sentence, paragraph, or modality. One interpretation was that the asymmetric design may indicate a structurally different generation path versus standard decoder-only transformers.
Local inference feasibility was discussed around the model’s reported 552B parameter scale, with commenters arguing it is impractical even for high-end local setups such as multiple DGX Spark/Strix-class systems. The suggested practical workflow was to use larger DeepSeek V4-class models for planning, then smaller/distilled models such as Q38-27B, Q38-35B-Distill, or Ornith35B for execution in local agentic coding pipelines.
A technically notable claim highlighted in the thread was a 437× KV-cache reduction since first generation, which commenters viewed as significant for long-context inference cost and memory scaling. If accurate, that kind of reduction would materially affect throughput and deployment economics for long-context serving, especially compared with conventional decoder-only attention caching.
Deepseek V4.1 Flash is 748B, not 552B (Activity: 575): OP inspected the Hugging Face safetensors and argues DeepSeek V4.1 Flash is ~748.5B parameters for backbone + engram—not 284B, 305B, 485B, or 522B—with a 551.566B backbone and 196.929B engram; including optional DSpark/MTP (14.225B) and vision encoder (0.485B) brings the stored model to ~763.21B params / 511.76 GB. The confusion is attributed to counting/metadata errors: e.g. an NVIDIA forum estimate undercounts the backbone, Hugging Face’s 485B likely miscounts FP4 packed weights as bytes rather than two params/byte, similar to GLM-5.3-Flash-NVFP4, and vLLM’s recipe inconsistently lists 522B before later correcting parameter details. The backbone is overwhelmingly MoE FFN experts: 543.582B params in FP4, with only ~7.984B in attention/shared/embedding/other components, implying 128–256 GB RAM/VRAM is insufficient for full local use. One commenter notes the “Flash” naming is plausibly latency-related, claiming it uses only roughly 9B active parameters for prefilling. Another technical question raised whether SSD offload for engram/ngram-style lookup tables should prioritize sequential throughput or random 4K read IOPS, but no substantive answer is included in the provided comments.
Commenters discussed that DeepSeek V4.1 Flash may report a much larger total size due to included n-gram/lookup-style components, but some argue these should not be counted like active neural parameters because they can be stored externally on SSD rather than loaded into VRAM/RAM as model weights.
A technical claim was made that the “Flash” variant is fast because it uses only around 9B parameters during prefill, implying the active compute path is far smaller than the headline 748B figure and may explain the latency-focused branding.
For local deployment, one commenter estimated that 256GB system RAM plus 64–96GB VRAM is sufficient, with the n-gram data hosted on any PCIe Gen 3+ NVMe SSD. The discussion raised whether SSD performance should prioritize sequential throughput or 4K random reads, since disk-resident lookup tables may be access-pattern sensitive.
Deepseek Has Soft Retired Deepseek V4 Pro (Activity: 1598): The image is a screenshot of a tweet saying DeepSeek is effectively “soft retiring” DeepSeek V4 Pro: V4 Pro traffic will be automatically routed to DS V4.1 Flash and billed at cheaper Flash pricing until V4.1 Pro launches. The stated rationale is that V4.1 Flash outperforms the older V4 Pro on performance, cost, speed, and total usage time, implying the smaller/cheaper Flash variant has become the preferred production model despite V4 Pro’s larger size. Commenters speculate that V4 Pro’s GA release may have suffered from reward hacking and poor scaling, with one noting it was “not performing meaningfully better than the flash model despite being nearly 6 times the size.” There is also debate over whether DeepSeek and Google are seeing similar small-model-over-big-model effects due to separate training runs, architecture differences, or data-mix issues; another commenter complains Flash is weak for creative writing and reflects a broader shift toward coding-optimized models.
Several commenters argued DeepSeek V4 Pro GA underperformed relative to its size, with one claiming it showed a “high degree of reward hacking” and was not meaningfully better than the Flash model despite being nearly 6× larger. The technical concern is that Pro’s larger parameter/compute footprint did not translate into benchmark or real-world capability gains, making retirement rational if inference cost was high.
A thread compared DeepSeek and Google cases where smaller “Flash” variants outperform or match larger models, suggesting these may not be simple distillations from one large training run. Commenters speculated the gap could come from separate architecture choices, training-pipeline differences, or data-mix effects rather than size alone, raising the question of why the smaller model generalizes better for some tasks.
Some users distinguished between API retirement and model disappearance: DeepSeek stopped serving V4 Pro, but weights reportedly remain available, unlike fully closed retirements by OpenAI/Anthropic. Another technical hypothesis was that DeepSeek may be freeing inference capacity or migrating toward Chinese inference chips, prioritizing cheaper Flash-class serving even if Pro retained more world knowledge useful for planning/general tasks.
DeepSeek-V4.1-Flash surprised .... (Activity: 537): The image is a reaction meme, but it highlights a technical claim that DeepSeek-V4.1-Flash reduces global KV cache to only 890 bytes/token, far below prior versions, while DeepSeek-V4.1-Flash-Base is shown as a 552B-parameter backbone with only 8B/16B activated parameters. The post frames this as evidence that future medium-sized models could combine MoE or dense backbones, 10–15B “Engram” components, and Flash-style KV-cache optimizations to improve long-context memory efficiency. Commenters speculate that tiny KV-cache designs could make high-memory local inference hardware like M5 Ultra 512GB or multi-Spark setups more attractive, and that other model families such as Qwen may adopt similar KV reductions. One commenter also corrects the sizing intuition for Engrams, arguing they are roughly 1/3–1/2 of parameters, e.g. a 30B dense backbone would pair with about a 10–15B Engram.
Commenters focused on memory pressure and hardware feasibility, noting that strong “AA scores” could make very-high-memory local inference setups like M5 Ultra 512GB and multi-Spark configurations more attractive. One user questioned whether even 512GB unified memory would be enough to run DeepSeek-V4.1-Flash “comfortably” when using multiple subagents, implying KV-cache and concurrency overhead may dominate beyond raw model weights.
A technical thread discussed architectural parameter allocation: engrams were estimated at roughly 1/3 to 1/2 of total parameters, so a 30B dense backbone would imply an additional 10B–15B engram component, for about 40B–45B total parameters. Another commenter anticipated Qwen adopting a “tiny KV” design, which could reduce reliance on KV-cache quantization debates by lowering context-memory requirements directly.
NASA, as part of its continuing effort to reduce paperwork and respondent burden, under the Paperwork Reduction Act (PRA), invites the general public and other Federal agencies to take this opportunity to comment on proposed and/or continuing information collections.
The most consequential AI news of the past year came from a standards body. In December 2025, Anthropic donated the Model Context Protocol to the newly formed Agentic AI Foundation, a directed fund under the Linux Foundation cofounded by Anthropic, Block, and OpenAI, with support from Google, Microsoft, AWS, Cloudflare, and Bloomberg. Six months earlier, […]
During a service incident, a customer-remediation workflow is moved onto an emergency route because the situation is critical and the team needs a fast resolution. Approvals are shortened, a priority queue is opened, and an on-call agent is cleared to use an alternate procedure until the service recovers. The incident ends, but the route stays […]
Perplexity uses Astra to write communications, change software, and monitor production systems, and checks in much less frequently than with earlier models.