Context Window is the daily AI brief for applied AI builders. Subscribe free →

Monday, July 27, 2026

Jul 28Jul 26

Feature

Jensen Huang Criticizes Closed AI for Hindering Forensic Efforts

Open-weight models proved crucial in diagnosing the Hugging Face incident, highlighting a debate on AI transparency.

Why it mattersAI infrastructure is at a crossroads as incidents like the Hugging Face breach highlight the need for transparency. For AI developers, balancing open and proprietary models is crucial to ensuring security without stifling innovation.

Read full article →

Sign up for the daily AI brief

Each morning: models, tools, research, and conversations distilled from across the web — with why they matter for builders.

Around the Web

AI Bubble Trouble

OpenAI Drama

AI Safety Concerns

Honorable mentions

Kimi-K3 is published on HuggingFace

Moonshot's latest model Kimi-K3 is available on HuggingFace since today. And it's another good news for open-weight AI and for the future of open-source AI It's a 2.8T-parameters Moonshot's SOTA model with 1 million tok…

r/artificial

Tools

Graphify-Labs/graphify

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local determi…

github | stars 96,987

JuliusBrussee/caveman

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

github | stars 93,474

DietrichGebert/ponytail

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

github | stars 90,261

zai-org/GLM-5.2

Trending AI model on Hugging Face — text-generation.

🤗huggingface

nexu-io/open-design

🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images &…

github | stars 81,911

santifer/career-ops

Open-source AI job search: scan job portals, evaluate listings with a structured A-F rubric into a 1.0-5.0 score, tailor your CV, track applications — runs locally in your AI codi…

github | stars 61,825

Research

SceneActBench: Can Agents Act on the 3D Scenes They See?

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete…

arxiv

Playbooks

A Field Guide to Rapidly Improving AI Products

Most AI teams focus on the wrong things. Here’s a common scene from my consulting work: AI TEAM Here’s our agent architecture – we’ve got RAG here, a router there, and we’re using this new framework for… ME [Holding up…

hamel

Demystifying evals for AI agents

The capabilities that make agents useful also make them difficult to evaluate. The strategies that work across deployments combine techniques to match the complexity of the systems they measure. \n

anthropic

The Microsoft Agent Framework Harness is now released

Your agents can now be built on a stable, batteries-included harness – the loop, planning, memory, context management, approvals, and telemetry that turn a model into an agent that actually does things – in both Python…

devblogs.microsoft

Evals Skills for Coding Agents

Today, I’m publishing evals-skills , a set of skills for AI product evals 1 . They guard against common mistakes I’ve seen helping 50+ companies and teaching 4,000+ students in our course . Why Skills for Evals Coding a…

hamel

Meet your agent harness and claw

Part 1 of Build your own claw and agent harness with Microsoft Agent Framework. In the overview we said a “claw” is really just an agent harness: a loop around a model, wired up with tools, planning, memory, and more. I…

devblogs.microsoft

News

Hetzner is working on LLM Inference

Hetzner has launched an experimental LLM inference API. I tested its Qwen model—and have a few guesses about where the product could go next.

sliplane

Quoting Boris Cherny

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is v…

simonwillison

Samsung AI Smart Glasses Prepare for a Fall 2026 Release

Samsung’s latest unveiling at the Galaxy Unpacked event has introduced the world to its AI-powered smart glasses, a product built on the Android XR platform. These glasses integrate advanced features like real-time tran…

geeky-gadgets