Context Window is the daily AI brief for applied AI builders. Subscribe free →

Saturday, July 25, 2026

Jul 26Jul 24

Feature

Curated Claude Code: A New Agent Harness with Intake Gate

Developers can now manage AI agent workflows more flexibly with YAML-defined orchestration.

Read full article →

Sign up for the daily AI brief

Each morning: models, tools, research, and conversations distilled from across the web — with why they matter for builders.

Around the Web

OpenAI Drama Alert

Is Open AI API down?

Hi everyone, My app hsas started throwing error which was calling Open AI api. When I am trying to login to https://developers.openai.com/ , it is throwing error. I have tried all my google accounts, is it happening wit…

r/OpenAI

AI in Daily Life

Honorable mentions

Tools

ultraworkers/claw-code

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

github | stars 194,899

Graphify-Labs/graphify

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local determi…

github | stars 95,660

JuliusBrussee/caveman

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

github | stars 92,909

DietrichGebert/ponytail

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

github | stars 89,266

nexu-io/open-design

🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images &…

github | stars 81,474

zai-org/GLM-5.2

Trending AI model on Hugging Face — text-generation.

🤗huggingface

Research

White Box Evidence Packages for Policy Audit Reports

As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that ques…

arxiv

Diffusion Language Model for Recommendation

Large language model (LLM)-empowered recommender systems have emerged as a promising paradigm for generative recommendation, leveraging their strong semantic reasoning and generative capacity to model complex, diverse u…

arxiv

Playbooks

A Field Guide to Rapidly Improving AI Products

Most AI teams focus on the wrong things. Here’s a common scene from my consulting work: AI TEAM Here’s our agent architecture – we’ve got RAG here, a router there, and we’re using this new framework for… ME [Holding up…

hamel

Demystifying evals for AI agents

The capabilities that make agents useful also make them difficult to evaluate. The strategies that work across deployments combine techniques to match the complexity of the systems they measure. \n

anthropic

The Microsoft Agent Framework Harness is now released

Your agents can now be built on a stable, batteries-included harness – the loop, planning, memory, context management, approvals, and telemetry that turn a model into an agent that actually does things – in both Python…

devblogs.microsoft

Evals Skills for Coding Agents

Today, I’m publishing evals-skills , a set of skills for AI product evals 1 . They guard against common mistakes I’ve seen helping 50+ companies and teaching 4,000+ students in our course . Why Skills for Evals Coding a…

hamel

Meet your agent harness and claw

Part 1 of Build your own claw and agent harness with Microsoft Agent Framework. In the overview we said a “claw” is really just an agent harness: a loop around a model, wired up with tools, planning, memory, and more. I…

devblogs.microsoft

How We Benchmark Deep Agents

We revamped how we benchmark Deep Agents. Here's the eval setup we run in Harbor across coding, conversation, and retrieval, and how we use it to ship changes.

langchain

News

Quoting Boris Cherny

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is v…

simonwillison

Hetzner is working on LLM Inference

Hetzner has launched an experimental LLM inference API. I tested its Qwen model—and have a few guesses about where the product could go next.

sliplane