❌

Reading view

Estimating from No Data: Deriving a Continuous Score from Categories

A walkthrough of and the maths behind using low-capacity networks to acquire fine-grained scoring when only categorical labelling is available for training

The post Estimating from No Data: Deriving a Continuous Score from Categories appeared first on Towards Data Science.

  •  

Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG

Enterprise Document Intelligence [Vol.1 #7sexies] - The unit of retrieval doesn’t have to be a page or a paragraph. When the corpus carries tables, each body row with its column headers is a chunk in its own right, and it’s often the one row the reader asked about

The post Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG appeared first on Towards Data Science.

  •  

The Types of Dimensions in a Star Schema, and How to Use Them

Dimensions are one of the two main object types in dimensional modelling. But what are the different types of dimensions? And how can you use them?

The post The Types of Dimensions in a Star Schema, and How to Use Them appeared first on Towards Data Science.

  •  

Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions

AI systems should not automate a decision simply because they can provide a prediction. A decision system should consider how uncertain the prediction is and defer if a mistake would be costly.

The post Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions appeared first on Towards Data Science.

  •  

How Benders Decomposition Works, Part II: Feasibility Cuts

Learning about Farkas' lemma and how it can inform Benders decomposition to learn from infeasibility, applied to the capacitated facility location problem.

The post How Benders Decomposition Works, Part II: Feasibility Cuts appeared first on Towards Data Science.

  •  

Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One

Enterprise Document Intelligence [Vol.1 #14A] - Three questions tell you which shape a document collection has, and each shape wants a different architecture

The post Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One appeared first on Towards Data Science.

  •  

Making the Knowledge Layer a Graph You Actually Traverse

Why retrieval quality should be a property of the system, not of the question's wording? Rebuilding knowledge layer with graph traversal on every query, bitemporal edges, and two-threshold entity resolution.

The post Making the Knowledge Layer a Graph You Actually Traverse appeared first on Towards Data Science.

  •  

How to Scale an Integration Pipeline Without Breaking Correctness

A production account of scaling an enterprise integration pipeline from 500 to 8,000 events per second, and the two correctness guarantees the throughput work was never allowed to trade away.

The post How to Scale an Integration Pipeline Without Breaking Correctness appeared first on Towards Data Science.

  •  

Kimi K3’s 1M Token Context Window vs. RAG: Cost, Latency and Answer Quality

A controlled comparison of a top-5 RAG pipeline and a full 127,000 token prompt on the same 12 questions, same system prompt and same model. Graded blind on correctness, completeness and grounding.

The post Kimi K3’s 1M Token Context Window vs. RAG: Cost, Latency and Answer Quality appeared first on Towards Data Science.

  •  

Building Enterprise Agent Systems that People can Trust, Verify andΒ Improve

5 principles that determine whether an agent system succeeds in production, explained through one I built for a $100M+Β company.

The post Building Enterprise Agent Systems that People can Trust, Verify andΒ Improve appeared first on Towards Data Science.

  •  

Graph Engineering Isn’t About More Connections β€” It’s About Which Ones Get Used

Adding more communication pathways between agents doesn’t necessarily improve multi-agent performance. In a controlled, reproducible experiment across 50 runs, recovery remained remarkably stable from 20% to 100% relationship density. But as the network became denser, the fraction of edges actually used fell sharplyβ€”revealing a gap between configured connectivity and behavioral connectivity.

The post Graph Engineering Isn’t About More Connections β€” It’s About Which Ones Get Used appeared first on Towards Data Science.

  •  

Webwright: Why AI Web Agents Should Write Code, Not Click

For years, web agents have worked one click at a timeβ€”and often fallen apart on long tasks. Microsoft Research’s Webwright makes a different bet: give the model a terminal and let it write the program instead. On long-horizon tasks, the same GPT-5.4 model jumps from 33.5% to 60.1% success. And instead of leaving behind a click trace, it leaves something you can actually use again: a command-line tool.

The post Webwright: Why AI Web Agents Should Write Code, Not Click appeared first on Towards Data Science.

  •  
❌