
Key takeaways
- A skill should reveal when it activates, what context it loads, and what job it owns.
- Skill prose does not replace runtime permissions, network limits, or independent review controls.
- Trust comes from realistic tests, visible state, explicit success criteria, and declared recovery behavior.
What is a context engineering skill?
A context engineering skill is a task-specific agent skill for shaping what an agent sees and how it approaches a job. Context engineering curates the information entering an LLM's limited attention at each step. The practical goal is not to load everything that might be relevant. It is to find the smallest set of high-signal tokens that makes the desired outcome more likely.
A skill turns part of that context strategy into something inspectable. In Claude, skills are task-specific and load when relevant. The main body of a Claude SKILL.md holds procedural knowledge such as workflows, guidance, and best practices. Claude Code can also run a command and place its output into the content before the model sees it. That makes the loaded context potentially dynamic, not just a static prompt.
| Reader intent | Object being inspected | Concrete example | Question it can answer |
|---|---|---|---|
| Understand the discipline | Context engineering | The information admitted to the model's attention at each step | What should the agent see now? |
| Inspect platform instructions | A Claude SKILL.md | Procedural workflows, guidance, and best practices | What procedure loads for this task? |
| Browse possible skill areas | A public collection | Murat Can Koylan's collection covers context engineering, multi-agent architectures, and production agent systems | Which kinds of work have been packaged as skills? |
| Break down implementation work | A planning-and-task-breakdown skill | A skill that turns a specification into small, verifiable tasks with acceptance criteria and dependency order | Can this specification become an ordered task plan? |
That map also answers the question, "What are the most useful agent skills?" Usefulness starts with a defined job, not a broad label. A planning skill is useful when the work is to decompose a specification. A repository collection is useful for discovery, but the collection itself does not answer whether one package has clear activation, authority, tests, state, or recovery. Those boundaries require a closer inspection.
What makes a context engineering skill useful in practice?
We use six dimensions to inspect a skill as a delegation contract. The rubric is not a performance score. It exposes what the package declares, what the surrounding system must control, and what still needs a test.
| Dimension | Repository evidence to inspect | Concrete test | Resulting decision |
|---|---|---|---|
| Activation | Invocation settings and trigger description | In Codex, set allow_implicit_invocation to false and confirm that prompt-based invocation stops while explicit $skill invocation remains | Decide whether activation should be implicit, explicit, or both |
| Context | The order of system context, task instructions, examples, input data, and output format | Inspect one run and confirm that those five parts appear in that order | Decide whether the agent receives the intended material in a legible sequence |
| Authority | Tool and permission configuration outside the prose | Use Claude Code permissions to specify what the agent may and may not do | Decide whether the runtime authority matches the job |
| Acceptance | A definition of completion tied to the execution trace | Extract the task and outcome from each trace step, then determine whether the outcome satisfied the task | Decide what counts as complete |
| Retained state | A stable thread identifier and a documented resume path | For a LangGraph workflow, interrupt it, store its state, and resume from the interrupt point | Decide whether interrupted work can continue from retained state |
| Recovery | Declared behavior for uncertainty and missing information | Exercise the documented fallback instructions, confidence scores, alternative interpretations, and missing-information signal | Decide whether failure behavior is explicit enough for the job |
The rubric makes a polished description less persuasive on its own. A skill may explain a strong procedure yet leave activation implicit. It may include a resume mechanism while never defining acceptance. Treat each blank cell as an unanswered engineering question, not as proof that the skill is unsafe or ineffective.
How should you inspect a skill repository?
Inspect a context-engineering skill GitHub repository from the outside in. First identify the one job the skill owns. Codex guidance says each skill should remain focused on one job, and its steps should use imperative language with explicit inputs and outputs. A description that names the topic but not the task gives you little help with activation or evaluation.

Then work through this teardown template:
- Job scope. Write the job in one sentence. If the package contains several distinct jobs, record each one instead of hiding them under a broad skill name.
- Trigger description. Find what the skill does and when it should trigger. The Viget Agent Skills contribution rules require both in a contribution description.
- Inputs and outputs. Locate imperative steps that name what each step receives and produces. This turns a general method into a procedure you can inspect.
- Script justification. Ask why each script exists. Codex skill authors are advised to prefer instructions unless deterministic behavior or external tooling is required.
- Dependency record. For every script, find its runtime and package requirements. Viget's repository requires skill scripts to be self-contained and their dependencies documented.
- Prompt-trigger test record. Look for prompts tested against the description and evidence that the intended trigger behavior occurred. Viget also requires a contributed skill to be tested with at least one agent.
This is a teardown, not a popularity contest. Stars and forks do not tell you whether a trigger is too broad, an input is missing, or a script relies on an undocumented dependency. The strongest repository evidence is close to the behavior you need to trust: the description, the instructions, the executable material, its dependencies, and the test record.
Scripts deserve particular care because they change the package from guidance into executable work. Their presence is not automatically a flaw. Deterministic behavior or external tooling can justify them. But that justification should be visible enough that a maintainer can understand why prose alone was insufficient and what the script needs in order to run.
Trigger testing deserves the same precision. A description can sound focused while matching prompts that should not activate it. It can also fail to match the task it was written to handle. Test prompts against the description and record the observed trigger behavior. If invocation must remain deliberate in Codex, explicit $skill invocation can stay available while implicit invocation is disabled.
The completed template should leave you with a small set of concrete decisions: whether the job is narrow enough, whether activation is predictable, whether inputs and outputs are legible, whether executable parts are justified, whether dependencies are known, and whether trigger behavior has actually been exercised. It still cannot tell you what the runtime permits. That boundary sits elsewhere.
Make the teardown maintainable
A useful teardown should survive the first review. Save the job statement beside the trigger prompts you exercised. Pair each script with the reason it exists and the dependencies it needs. Keep the observed output with the prompt that produced it. This creates a compact record that a maintainer can revisit when the description, instructions, or executable material changes.
Do not compress that record into a single pass or fail label. One agent test can satisfy a repository contribution rule, but the teardown still has six separate questions. A skill might have a clear trigger and poor input definitions. Another might use precise steps but leave a script dependency undocumented. Keeping those findings separate shows where maintenance work belongs.
The template is equally useful when authoring a new package. Start with the one-job sentence, write the trigger description, and state the input and output of each imperative step. Add a script only when deterministic behavior or external tooling requires one. Then document its dependencies and test prompts against the description. This order keeps the visible contract close to the behavior you are asking an agent to perform.
Runtime restrictions remain outside skill prose
A skill can describe the work, but its text is not the whole harness. Repository-level instruction files keep Codex aware of project norms while Codex continues to inherit global defaults. That inheritance is useful for carrying local conventions into a project, but it is separate from the controls that constrain execution.

Network authority is one clear example. Codex command network access can be limited with destination rules, and those rules apply to scripts, programs, and subprocesses. If a skill launches a script that starts another process, the network boundary still belongs to the runtime configuration. It does not become a property of the skill's prose.
Review instructions are also not a substitute for independent filesystem and network restrictions. A reviewer can assess behavior within a process, while the surrounding system enforces what resources the process can reach. Keeping those mechanisms distinct makes the control surface easier to inspect: project instructions carry norms, destination rules constrain network access, and independent restrictions protect filesystem and network boundaries.
When reviewing a repository, trace each control to the layer that owns it. Read repository instructions for project norms, then account for the global defaults Codex continues to inherit. Inspect destination rules for the network access available to commands and their subprocesses. Finally, identify the independent filesystem and network restrictions. Keep reviewer instructions separate from those enforced boundaries.
This layered view also helps when a skill contains scripts. The teardown can establish why a script exists and which dependencies it declares. Runtime inspection answers where that script and any subprocesses may connect. Neither view replaces the other, so adoption should not stop after reading the main instruction file.
This separation is central to harness engineering. The skill carries task knowledge. The harness carries the execution paths and controls around that knowledge. You need both views before trusting a package with consequential work.
How do you test whether a skill works?
Start with a realistic failure, then turn it into an evaluation task. A failed step in an agentic system can send the agent onto a different trajectory and produce unpredictable later outcomes. That makes a clean final answer an incomplete test. The transcript matters because it shows the path the agent took.
Consider a planning skill that receives a difficult specification but produces tasks without acceptance criteria or dependency order. Use that observed failure as the evaluation case. The task is the original specification. Success requires small, verifiable tasks, explicit acceptance criteria, and dependency ordering. The grader checks those named requirements, while transcript review checks how the agent interpreted the specification and where the plan lost structure.
Keep two results separate. Plan Quality evaluates whether a proposed plan is complete, realistic, and efficient. Plan Adherence evaluates whether execution follows the plan, user constraints, or expected workflow. These are vendor-defined metrics, not platform-neutral standards, but the distinction is useful: one result examines the plan itself, while the other examines execution against it.
A sound evaluation set uses realistic tasks, clear success criteria, carefully designed graders, and problems difficult enough to reveal meaningful behavior. Improve the signal over repeated runs, and review transcripts instead of relying only on aggregate results. Early evaluation work can turn failures into test cases, use those cases to prevent regressions, and replace guesswork with metrics.
Build the test from the failure you observed
The most useful first test is often already in front of you. Preserve the failed task, state the outcome that should have satisfied it, and turn those requirements into unambiguous success criteria. Choose a grader that can judge those criteria, then inspect the transcript for the step where the trajectory changed. If the task is too easy to expose the failure again, use a sufficiently difficult problem that exercises the same behavior.
As the evaluation develops, improve its signal instead of multiplying vague examples. A realistic task with a clear grader gives you a result you can act on. Transcript review gives you the path behind that result. Together, they let the original failure become a regression case instead of an anecdote that disappears after the prompt is edited.
Do not report the six Skill Contract dimensions as a single score unless each score has an explicit test and result. A repository can pass a trigger test and still lack a credible recovery path. It can produce a good plan and still violate the expected workflow during execution. Keep the results attached to the behavior they measure. For a related view of retained evidence, see our work on an AI audit trail.
When should an agent plan or pause?
Ask the agent to plan before implementation when a coding task is complex, ambiguous, or difficult to describe. Codex recommends planning first for those tasks. The plan gives you an object to inspect before execution begins, and it creates something that Plan Quality and Plan Adherence can evaluate separately.
Pause at a different boundary: before a flagged irreversible or sensitive operation. A high-risk action gate stops that operation until a reviewer explicitly approves the next step. The gate is not a general request for human attention. It is attached to the action whose consequence warrants approval.
Use two checks when authoring or adopting a skill:
- Does this task need a plan because it is complex, ambiguous, or difficult to describe?
- Does any flagged irreversible or sensitive operation require explicit reviewer approval before it continues?
These checks put human judgment where it can change the course of work. They also complete the contract: activation starts the skill, context shapes the work, authority limits action, acceptance defines completion, retained state supports continuation, and recovery covers uncertainty or failure.
Planning and pausing should remain two explicit decisions. Requesting a plan creates a proposed route through difficult coding work. A high-risk gate stops a flagged operation and waits for approval. Review the plan before implementation when the task meets the complexity, ambiguity, or description test. Require approval when the operation is flagged as irreversible or sensitive.
Use the Skill Contract rubric on the next repository you evaluate. Record the evidence and the test result for each dimension, then carry the unanswered runtime questions into the surrounding harness.