Context Window is the daily AI brief for applied AI builders. Subscribe free →

Sunday, July 26, 2026

Jul 27Jul 25

Feature

Running a 28.9M Parameter LLM on an $8 Microcontroller

Exploring the practical implications and challenges of deploying large language models on low-cost hardware.

Why it mattersAI infrastructure is evolving as LLMs can now run on low-cost microcontrollers, opening up new possibilities for edge computing. This shift challenges developers to balance accessibility and affordability with performance and security concerns in resource-constrained environments.

Read full article →

Sign up for the daily AI brief

Each morning: models, tools, research, and conversations distilled from across the web — with why they matter for builders.

Around the Web

Cost-Cutting Inference Tricks

AI Models on a Budget

Kimi K3 gets open weighted tomorrow!

Kimi K3 is supposed to get open weighted tomorrow! Can't run it or even a model a hundred times smaller lol, but its still a great win for open source. For me, personally im more awaited for the new inference providers…

r/LocalLLaMA

AI Regulation Shenanigans

LLM Benchmarking Fun

We compared different LLMs on IMO 2026 [R]

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math problems are quite…

r/MachineLearning

Honorable mentions

Tools

Graphify-Labs/graphify

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local determi…

github | stars 96,251

JuliusBrussee/caveman

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

github | stars 93,135

DietrichGebert/ponytail

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

github | stars 89,700

nexu-io/open-design

🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images &…

github | stars 81,675

zai-org/GLM-5.2

Trending AI model on Hugging Face — text-generation.

🤗huggingface

santifer/career-ops

Open-source AI job search: scan job portals, evaluate listings with a structured A-F rubric into a 1.0-5.0 score, tailor your CV, track applications — runs locally in your AI codi…

github | stars 61,643

Research

White Box Evidence Packages for Policy Audit Reports

As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that ques…

arxiv

Playbooks

A Field Guide to Rapidly Improving AI Products

Most AI teams focus on the wrong things. Here’s a common scene from my consulting work: AI TEAM Here’s our agent architecture – we’ve got RAG here, a router there, and we’re using this new framework for… ME [Holding up…

hamel

Demystifying evals for AI agents

The capabilities that make agents useful also make them difficult to evaluate. The strategies that work across deployments combine techniques to match the complexity of the systems they measure. \n

anthropic

The Microsoft Agent Framework Harness is now released

Your agents can now be built on a stable, batteries-included harness – the loop, planning, memory, context management, approvals, and telemetry that turn a model into an agent that actually does things – in both Python…

devblogs.microsoft

Evals Skills for Coding Agents

Today, I’m publishing evals-skills , a set of skills for AI product evals 1 . They guard against common mistakes I’ve seen helping 50+ companies and teaching 4,000+ students in our course . Why Skills for Evals Coding a…

hamel

Meet your agent harness and claw

Part 1 of Build your own claw and agent harness with Microsoft Agent Framework. In the overview we said a “claw” is really just an agent harness: a loop around a model, wired up with tools, planning, memory, and more. I…

devblogs.microsoft

News

Quoting Boris Cherny

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is v…

simonwillison

KDnuggets Weekly Roundup: Week of July 20, 2026

Top 5 MCP Servers for High Performance Agentic Development • 10 Newsletters Keeping You Ahead in AI • Kaggle + Google’s Free 5-Day Agentic AI Course • Language Model Hallucination Evaluation with GraphEval

kdnuggets

Hetzner is working on LLM Inference

Hetzner has launched an experimental LLM inference API. I tested its Qwen model—and have a few guesses about where the product could go next.

sliplane