When your AI agent gives a wrong answer, the fix is rarely a better model. More often the agent was missing a document, had 40 pages of irrelevant history in its window, or picked the wrong tool because two tools had near-identical descriptions. The work of fixing that has a name now. In June 2025, Shopify CEO Tobi Lutke wrote that he preferred the term "context engineering" over prompt engineering, describing it as "the art of providing all the context for the task to be plausibly solvable by the LLM." Andrej Karpathy added his "+1" a few days later. The term stuck, and startups building agents now hire for it.
Key Takeaways
- Context engineering is the design of everything a model sees at each step: instructions, retrieved data, tools, memory, and history.
- It's usually a skill inside an AI engineering role rather than a separate job title. Hire the skill, whatever you call the role.
- Screen with a real failing agent. Ask the candidate to diagnose why it fails and fix the context, not the model.
- Look for evaluation habits. A context engineer who can't measure whether a change helped is guessing.
What Context Engineering Actually Is
LangChain's Harrison Chase defined it in June 2025 as "building dynamic systems to provide the right information and tools in the right format such that the LLM can plausibly accomplish the task." Anthropic's Applied AI team, in Effective context engineering for AI agents, put the goal plainly: "find the smallest set of high-signal tokens that maximize the likelihood of your desired outcome."
In practice, the work falls into a few buckets. LangChain's follow-up post groups them as write, select, compress, and isolate:
| Strategy | What it means | Example task |
|---|---|---|
| Write | Save information outside the window for later | Agent keeps a notes file across a long task |
| Select | Pull in only what's relevant | Retrieval that returns 5 good chunks, not 50 loose ones |
| Compress | Shrink what's already there | Summarize old conversation turns before the window fills |
| Isolate | Split context across agents or sandboxes | A sub-agent researches, returns a short summary |
Anthropic's post adds techniques like compaction, structured note-taking, just-in-time retrieval, and writing system prompts at the "right altitude": specific enough to guide the model, general enough to handle cases you didn't predict.
Do You Need a Context Engineer or an AI Engineer?
Usually you need an AI engineer or an AI agent developer who is strong at context engineering. The title "context engineer" is new, and few people have held it. Your candidates will more often call themselves AI engineers, applied ML engineers, or backend engineers who've built LLM features for the last two years.
A dedicated context-focused hire makes sense when you already have an agent in production, it mostly works, and the remaining failures come from what it sees rather than what it can do. At that stage, someone who spends all day on retrieval quality, tool design, memory, and evaluation can move your success rate more than another feature developer.
The Skills to Screen For
Based on the sources above, a strong candidate can show real work in:
- Retrieval design. Chunking, ranking, and filtering so the model gets relevant data and little else.
- Tool design. Clear tool names and descriptions, non-overlapping tools, and outputs the model can use.
- Memory. What an agent should remember between steps and between sessions, and what it should forget.
- Compression. Summarizing and trimming history without losing the facts that matter.
- Multi-agent isolation. When to hand a subtask to a sub-agent with a clean window.
- Evaluation. Test sets and traces that show which failures come from missing or badly formatted context.
The last one separates senior people from hobbyists. Anyone can tweak a prompt until one demo works. A context engineer builds a set of 50 or 100 real cases and proves the change helped across all of them. How to vet an engineer's AI skills covers more general screening.
A Practical Exercise
Give the candidate a small agent that fails in known ways. For example, a support agent with access to a help-center search tool and an order-lookup tool, plus 20 test questions it gets wrong about a third of the time. Give them three hours.
Don't let them change the model. Ask them to find out why it fails and fix the context. Strong candidates read traces before touching anything. They'll notice the search tool returns whole articles when a paragraph would do, that the two tool descriptions overlap, or that the system prompt contradicts itself. Then they re-run the test set and show you the numbers before and after.
A Concrete Version
A 10-person startup sells an AI assistant that answers questions about a customer's contracts. In testing it scores well. With real customers, who upload 200-page agreements, answers get vague and it sometimes cites the wrong clause.
Their new hire, a senior engineer with strong context skills, starts by pulling 80 real failed questions into an evaluation set. The baseline gets 52 of them right. She changes three things: retrieval returns clause-level chunks with section numbers instead of full pages, the agent writes a short list of relevant clause numbers to a notes file before answering, and conversation history older than five turns gets summarized. The same set now scores 71 of 80. The model never changed. The team ships the fix in three weeks and keeps the evaluation set as a regression test.
The Honest Counterpoint
Context engineering is a real skill with a trendy name, and trendy names attract resume padding. Plenty of candidates will list it after reading a few blog posts. The exercise above filters most of them out.
Don't hire for it too early, either. If you haven't shipped an AI feature yet, your problem is product and basic engineering, and a strong generalist AI engineer will do more. The specialist pays off once you have real traffic, real failures, and an evaluation set to improve against.
Frequently Asked Questions
Is context engineering the same as prompt engineering?
Prompt engineering is a part of it. Context engineering also covers retrieval, tools, memory, and history, which is most of what a production agent sees.
What background do good context engineers come from?
Often backend or search engineering, sometimes ML. Experience with retrieval systems and a habit of measuring results matter more than an AI title.
Can a strong backend engineer learn this?
Yes, if they're curious and rigorous about evaluation. Pair them with your existing AI work and give them a real failure set to improve.
The Bottom Line
Hire for context engineering as a skill: retrieval, tools, memory, compression, and above all evaluation. Test candidates on a failing agent, not a whiteboard. Ruzora sends a vetted shortlist of senior AI engineers within 72 hours. Hire AI engineers in LATAM, or request a shortlist for your agent team.
Roberto Espinoza is CEO of Ruzora, which helps US startups hire pre-vetted senior LATAM engineers, with a vetted shortlist in 72 hours. See available engineers.
