❌

Normal view

Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge

3 June 2026 at 17:15
Google DeepMind’s Gemma 4 12B model brings agentic, multimodal AI capabilities to everyday laptops with 16GB of RAM, enabling local data processing and visual insight generation. Users can leverage this model on macOS through the Google AI Edge Gallery for dynamic Python code execution and visualization, as well as via Google AI Edge Eloquent for completely offline voice dictation and text editing. Additionally, developer workflows are enhanced by the LiteRT-LM CLI's new serve command, which creates an industry-compatible local endpoint to power fully-local AI tools and agents.

Ground truth is a process, not a dataset

3 June 2026 at 15:56
Today, the key challenge in AI isn’t only how to build better models; it’s how to build evaluation systems that can keep up. Search-augmented AI systems can now produce deep research reports — long, polished syntheses of many sources that increasingly resemble expert analysis. But those reports are useful only if their claims are supported by the underlying literature. Most existing fact-checking tools work best when a claim can be matched to a short quote or a single document. But in AI-generated research reports, a single sentence may combine evidence from several sources. It can depend on the surrounding report for context, and it might compare assertions in a way that no single source does on its own. When Amazon’s Artificial General Intelligence (AGI) group started working on the problem of evaluating AI-generated research reports, we thought that the main technical challenge would be building a stronger AI fact checker. But before you can evaluate an AI fact checker, you need a benchmark, a standardized test set used to measure performance. And in this setting, building the benchmark turned out to be at least as hard as building the model. Traditionally, we view the ground truth for a problem as a fixed dataset. But we discovered that to evaluate complex AI properly, ground truth has to become a process. We call that process audit-then-score, and we present it, together with two accompanying datasets, in a paper we recently published to arXiv. When static datasets break down In the standard method for measuring AI performance, human experts label examples, those labels become the “ground truth” (the undisputed correct answers), and models are scored against them. To test this approach with AI-generated research reports, we recruited PhD-level specialists from fields such as computer science, control theory, education, public health, and environmental engineering. We asked them to verify claims from reports in their own specialties, mixing in a hidden set of claims whose answers we already knew. The result was sobering. In a controlled study, unassisted experts achieved only 60.8% accuracy on the hidden set of known answers. The issue was not a lack of expertise. It was that assessing deep-research factuality is an unusually demanding task. Verifying a single claim can require long-context reading, cross-document synthesis, and sustained attention. Normally, in machine learning, when a model disagrees with a benchmark, we assume the model made a mistake. But we realized that, in cognitively demanding tasks like deep research, disagreement should not automatically be treated as a model failure. Sometimes, a model’s “error” is actually a signal that the benchmark itself is ambiguous, incomplete, or wrong. Audit, then score Instead of treating the initial expert labels as unquestionable ground truth, we decided to use the models to actively scrutinize the benchmark. This is the core idea behind the audit-then-score protocol. Our paper introduces the protocol alongside DeepFact-Bench, a shared test set for comparing systems, and DeepFact-Eval, a system that checks whether literature supports report claims. Here is how the protocol works: When our AI fact checker disagrees with the current benchmark answer, it is not simply penalized. Instead, it acts as a challenger and must submit concrete evidence and a written rationale for why it thinks the original human answer is wrong. An auditor — which can be a human expert — then steps in. Crucially, auditors do not start from scratch; they compare the challenger’s new evidence directly against the benchmark’s original rationale. If the challenger makes the stronger case, we revise the benchmark before we score the model. DeepFact-Eval reads the full report context, plans searches to cover the relevant literature, summarizes retrieved documents, and asks follow-up questions when key details are missing. It then produces both a verdict and a written explanation. This fundamentally changes what a benchmark is. A new role for human expertise One of the most striking things we found is that the same experts who were unreliable as one-shot labelers became far more reliable when placed in the role of auditor. Across four rounds of audit-then-score, accuracy on our hidden test set rose from 60.8% to 90.9%. When experts start from a blank page, they have to find the evidence, interpret it, and make a judgment on their own; when they audit a disputed claim, they can focus on comparing two concrete cases. This shift had significant impact. On DeepFact-Bench, DeepFact-Eval reached 83.4% accuracy when we used GPT-4.1 as the underlying model. That was higher than the 58.5% of the best traditional fact-checking system we tested and the 69.1% of a strong prior deep-research system. Evaluation as an evolving infrastructure This shift has implications beyond one paper or one task. If AI systems continue improving, to the point that they exhibit humanlike expertise, the community will increasingly run into settings where evaluation based on one-time human answers is not enough. In those settings, sustaining benchmark quality may require auditing, revision, calibration, and periodic revalidation. Evaluation will become an ongoing collaboration among humans, models, and the evidence they surface together. Acknowledgments: Yukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Markus Dreyer

Gemma 4 12B: The Developer Guide

3 June 2026 at 16:01
The newly released Gemma 4 12B is a dense, multimodal model designed for high-performance local AI execution on consumer devices. By introducing a novel, encoder-free architecture, it bypasses traditional visual and audio encoders to feed multimodal data directly into the LLM backbone.

Notice Announcing Fund for the Improvement of Postsecondary Education-Postsecondary Student Success Grants Program Competition

The Department of Labor and Department of Education (ED) announce the opportunity to apply for competitive grants for the Fiscal Year (FY) 2026 for the Fund for the Improvement of Postsecondary Education (FIPSE)--Postsecondary Student Success Grants (PSSG) Program, Assistance Listing Number (ALN) 84.116M.

Standards Participation and Representation Kudos (SPARK) Pilot Program

The United States Patent and Trademark Office (USPTO) is launching the Standards Participation and Representation Kudos (SPARK) Pilot Program to incentivize meaningful participation in standards development organizations (SDOs) by U.S. small and medium-sized businesses, universities, and non-profit organizations. Under the pilot program, examination of certain patent applications and ex parte appeals to the Patent Trial and Appeal Board (PTAB) may be expedited if the U.S.-domiciled juristic applicant meaningfully participated in a voluntary consensus-based SDO and meets the requirements specified in this notice. The application or appeal being expedited does not need to be related to the SDO participation. The expedited examination or appeal provides additional tangible value for the time and resources invested in standards development. Applications accepted into the pilot program for expedited examination will be advanced out of turn, that is, accorded special status, for examination until a first Office action is issued, and ex parte appeals accepted into the pilot program will be advanced out of turn before the PTAB. This notice sets forth the requirements of the pilot program and describes how the pilot program will be administered.

The OECD AI Policy Toolkit: Better AI policies for better lives

3 June 2026 at 06:54

Artificial intelligence (AI) is both a technology story and a policy challenge. Governments across sectors and regions are grappling with the same question: how to effectively support the safe, trustworthy development and use of AI in ways that align with their countries’ needs?

Whether setting a national AI strategy or designing concrete initiatives to implement it, governments need guidance that meets them where they are. From experience, I can attest that the hardest part is rarely agreeing on principles; it is finding concrete, comparable examples of how others made them work. That is the gap the OECD AI Policy Toolkit closes.

Released yesterday by the OECD under the Global Partnership on Artificial Intelligence (GPAI), the AI Policy Toolkit is the first version of a practical, non-prescriptive guide for policymakers to translate the OECD AI Principles into action—a deliberate shift from defining what good AI policy requires to showing how to build it.

What the Toolkit does

The Toolkit is an interactive, evolving platform to support policymakers throughout the AI policy cycle. It complements OECD.AI’s broader ecosystem of tools for data, analysis and AI governance.

The Toolkit helps governments target and prioritise where to act. Through AI-powered semantic search, it surfaces relevant policy examples and guidance drawn from real-world practice, turning the OECD’s accumulated evidence into options a policymaker can use the same day—rather than a library to be read.

Built with policy-makers, not just for them

A year ago, the 2025 OECD Ministerial Council Meeting set this work in motion. What followed was less a drafting exercise than a year of listening—and the Toolkit released today reflects what countries told us they needed.

Far from being a top-down exercise, the OECD Secretariat developed the Toolkit with end-users through co-creation across regions. Targeted interviews and four co-creation workshops across Southeast Asia, Latin America and Africa—one of which Costa Rica was proud to host—brought policymakers, industry and experts together to shape its design around how governments actually work and make decisions.

Not only did these co-creation workshops highlight both shared challenges and region-specific priorities. They grounded the Toolkit in fundamental policy questions:

  • How to navigate trade-offs between local and global AI models, or between innovation and regulation?
  • How to address infrastructure gaps, such as AI compute capacity?
  • How to scale AI in agriculture, education or healthcare?

Two lessons that shaped the Toolkit

Moreover, the collaborative approach to developing the Toolkit has yielded important collective lessons.

  • First, context is decisive: AI policy must reflect national needs and preferences, institutional capacity and levels of digital maturity.
  • Second, addressing shared global challenges such as managing risks posed by advanced AI systems or ensuring diverse linguistic and cultural representation in AI models requires international cooperation as well as tailored policy responses.

Our sincere thanks go to the governments and organisations that, alongside Costa Rica, made this possible—notably Italy, France, Korea, Japan, the United Kingdom, the European Union, the French Development Agency and the Inter-American Development Bank—and to the policymakers and experts who contributed their time and insight. I also commend the OECD Secretariat for its sustained work.

What comes next

The OECD Ministerial Council Meeting (MCM) marks the Toolkit’s first release, which is an important milestone, but it is far from the finish line.

As AI technologies and related policy issues develop, the OECD remains dedicated to ensuring the Toolkit stays relevant through regular updates by:

  • Refining and improving the Toolkit through ongoing feedback and iteration
  • Incorporating more policy examples and use cases to strengthen its practical relevance via the OECD.AI Policy Navigator
  • Expanding its coverage of emerging policy issues, including sector-specific guidance, infrastructure and regulatory approaches

From shared principles to shared practice

The OECD AI Policy Toolkit results from a collaborative effort to transform AI principles into implementation. By integrating OECD standards with regional insights, it guides policymakers in leveraging AI’s opportunities while responsibly and effectively managing its challenges.

The Toolkit’s success will be measured not by its launch but by the policies it helps shape. Its impact depends on sustained collaboration and support. A year from now, I expect us to point to concrete cases where this tool moved a country from principle to practice—better AI policies for better lives.

The post The OECD AI Policy Toolkit: Better AI policies for better lives appeared first on OECD.AI.

Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA

AI agents are changing how you interact with your PC. Creators, developers, and AI enthusiasts are already using these agents extensively to assist with...

AI agents are changing how you interact with your PC. Creators, developers, and AI enthusiasts are already using these agents extensively to assist with day-to-day tasks such as coding, video editing, and content management. NVIDIA and Microsoft are teaming up to enable the next generation of developers to build on-device agents on the Windows platform, with easier setup, native security…

Source

Deploy Self-Evolving Agents for Faster, More Secure Research with a Hermes Agent and NVIDIA NemoClaw

2 June 2026 at 16:00
Decorative image.AI agents are a powerful tool for synthesizing data to accelerate research, summarize information, and help teams make decisions faster. But combining internal...Decorative image.

AI agents are a powerful tool for synthesizing data to accelerate research, summarize information, and help teams make decisions faster. But combining internal data with public sources poses security challenges. This post shares an open source example using Hermes Agent with NVIDIA NemoClaw for product research across Outlook, Slack, and GitHub. NVIDIA OpenShell enforces a security-approved…

Source

❌