The map
The AI Engineering Skills Map is the spine of this deskAndrew Ng’s four skills — wrap, eval, ground, ship — are how this board is filed. If you are building an AI app, start here, not with a model card.
Latent Space5d agoEditorial
Setting type
Claude Computing
Pulling labs, repos, and the catch-up path…
Harnesses, graphs, loops, APIs you call this week — tagged with Ng’s application skills. Not a training desk. New here is the decoder. Starting a project this week? Catch up is the four-skill map — not more news.
The map
The AI Engineering Skills Map is the spine of this deskAndrew Ng’s four skills — wrap, eval, ground, ship — are how this board is filed. If you are building an AI app, start here, not with a model card.
Latent Space5d agoEditorial
The craft
Harness engineering is the new self-improvement storyLilian Weng’s map of recursive improvement is not sci-fi: it is loops, , and the wrapper around the model. If you only watch model cards, you will miss the job.
Lilian WengJul 4Editorial
Cookbooks for building AI apps: context, harness, graphs, cloud.
Eval
Trace the agent loop: OpenAI observability you ship withEval in the wrap: traces and integrations for agents you already call, not a training-lab dashboard.
OpenAIEditorial
Harness, loops, graphs, MCP, permissions, evals as acceptance.
Paper
WikiSkill: compile agent experience into persistent skillsA harness pattern, not a training run: turn traces into a wiki the next loop can actually reuse.
arXiv 2608.274542d agoEditorial
Paper
RedEvoAgent: red-team the execution harness, not the promptProduct-shaped red teaming: agents attacking the you ship, not unsafe-text on a chatbot.
Clone-to-ship: RAG, eval harnesses, agent wrappers. Not train-a-model.
Repo
deepset-ai/haystackPipelines for retrieval and eval, not a chatbot skin. How you ship grounding in production.
GitHub22h ago26k ★Live
Repo
promptfoo/promptfooWatchlist posts for people building AI apps. Originals, not architecture notes.
Course
Hugging Face agents course: ship an agent, freesmolagents, LlamaIndex, LangGraph, then a project with evals. Catch up already files the lab; this is the pointer.
Ren / X3d agoEditorial
X
An LLM cliché highlighter, 38 patterns inarXiv 2608.274392d agoEditorial
Stanford’s DSPy treats pipelines as code with optimizers. The anti-magic stance: if you cannot eval it, you cannot improve it.
stanfordnlp/dspyEditorial
Agent Development Kit is how Google wants you to build, eval, and deploy agents. Worth knowing even if you never ship on Vertex.
google/adk-pythonEditorial
Deep dive into self-improving evaluators in LangSmith, motivated by the rise of LLM-as-a-Judge evaluators plus research on few-shot learning and aligning human preferences.
LangChain4d agoLive
Eval harness for prompts and agents. Red-team and regression the wrapper, not the weights.
GitHub20m ago25k ★Live
A working eval for slop, not a vibe check.
Simon Willison / X2d agoLive
X
Stop testing models with a pelican on a bicycleEvals leaving toy SVG tests: long-horizon generation with a real budget.
Andrej Karpathy / XAug 2Live