Skipping permissions is the whole product question for coding agents. Read this as a security and UX harness, not a feature list.
Simon Willison2d agoLive
Setting type
Claude Computing
Pulling labs, repos, and the catch-up path…
Harnesses, graphs, loops, APIs you call this week — tagged with Ng’s application skills. Not a training desk. New here is the decoder. Starting a project this week? Catch up is the four-skill map — not more news.
Skipping permissions is the whole product question for coding agents. Read this as a security and UX harness, not a feature list.
Simon Willison2d agoLive
Cookbooks for building AI apps: context, harness, graphs, cloud.
Call it through Z.ai or Baseten — long-context wrap, MIT weights as a swap. Cookbook for the endpoint, not a training drop.
Z.ai4d agoEditorial
Context
Effective context engineering for AI agentsWindow, packing, routing: what belongs in context this turn. A lab you run on the wrap, not a model-card internals note.
Anthropic EngineeringEditorial
Wrap
Call Bedrock Converse in your app this weekWrap the AWS API — messages, tools, streaming — inside the product. Not a VPC lab.
AWSEditorial
MCP
Code execution with MCP: tools without stuffing the windowHarness lab: let the agent write code against MCP instead of packing every tool schema into context.
Anthropic EngineeringEditorial
Harness, loops, graphs, MCP, permissions, evals as acceptance.
Weng’s 2023 LLM Powered Autonomous Agents post is how a generation learned planning, memory, and tools. Read it before the new jargon.
Lilian WengJun 23Editorial
Stanford’s DSPy treats pipelines as code with optimizers. The anti-magic stance: if you cannot eval it, you cannot improve it.
stanfordnlp/dspyEditorial
Deep dive into self-improving evaluators in LangSmith, motivated by the rise of LLM-as-a-Judge evaluators plus research on few-shot learning and aligning human preferences.
LangChain4d agoLive
Clone-to-ship: RAG, eval harnesses, agent wrappers. Not train-a-model.
Repo
furkankly/zoetropeWatch a Claude Code session as a live flow graph. Read the harness, don’t guess what the agent just did.
GitHub5d ago663 ★Live
Watchlist posts for people building AI apps. Originals, not architecture notes.
X
Ramble into the window: voice as contextContext engineering by voice: more bits in the window without typing the spec.
Andrej Karpathy / XJul 21Live
Drew Breunig: model cost stopped papering over a weak harness. Context and routing are the job again.
Drew Breunig7d agoEditorial
Simon Willison’s CLI upgraded to OpenAI Python 3.x and httpx2. The adapter you call through, not a cookbook and not a new model.
Simon Willison8d agoEditorial
Course
Free company AI courses: Anthropic, Google, Meta, NVIDIA, Microsoft, OpenAIA Catch up pointer to vendor wrap/eval/ship courses. Not a second lab on this board.
Parag Pawar / XAug 7Editorial
Claude
Full Claude Course: practical Claude for shipping an appWrap Claude, add tools, ship a real project. Practical course, not a hustle bait.
Jobescape / XAug 11Editorial
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
OpenAIAug 13Live
X
Stop testing models with a pelican on a bicycleEvals leaving toy SVG tests: long-horizon generation with a real budget.
Andrej Karpathy / XAug 2Live
The Anthropic plugin for llm now tracks SDK 1.x. Upgrade the wrap layer this week so Claude calls keep compiling.
Simon Willison6d agoEditorial
Course
Become an AI engineer: LLM apps, APIs, context, tools, deployWrap, RAG, tool-use, evals, then ship. A six-month app-builder path, not a lab-training syllabus.
Rahul / XAug 11Editorial
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster.
OpenAIAug 13Live