Before You Build an AI Harness, Become the Harness
Before building software around an LLM, manually provide the tools and context it would have. The interventions reveal whether the workflow deserves a harness.
On this page
I am developing a coding harness for students working on projects that will be evaluated. The idea is to help them learn while they build: ask questions about their approach, examine what they are doing, correct misunderstandings, and provide the explanations they need to move forward.
That makes the goal more complicated than getting the code to run. If the system completes the project while the student understands very little of it, it has missed the point. I want a Socratic learning mechanism, but I also want it to provide useful guidance. Asking questions indefinitely would be another way to fail.
The project is still under development. I cannot claim that it improves learning yet. What I can describe is the rule I use to decide whether an idea deserves the effort of building a harness in the first place: before writing the software around an LLM, I try being that software myself.
I expose the information, execute the tools, return the results, and see whether the interaction can move toward the intended goal. It is an entry test. Before I invest in building the system, I want some evidence that there is a workable process to build around.
What I mean by an AI harness
For this article, an AI harness is the software around a language model that connects it to a working environment. It determines which tools are available, executes permitted actions, supplies relevant context, and manages the interaction as the task progresses. Depending on the application, it may also enforce limits, validate results, and hand decisions back to a person.
My test temporarily puts me in that position. The model communicates through chat. If it requests information, I retrieve it through a mechanism the eventual product could realistically offer. If it requests a tool operation, I execute it within the boundaries of the experiment and return the result. Then I see what it does next.
This is deliberately less elaborate than building the application. I can explore whether the workflow makes sense before committing to integrations, interfaces, and orchestration. However, the experiment only means something if I distinguish between enabling the model and quietly doing its work.
The easiest way to fool myself
Suppose the model asks for the wrong information. I understand the task, so I supply what it should have requested. Later, it misreads an error, and I explain the underlying problem. When it proposes an unsuitable next step, I steer it toward a better one. Eventually, we reach a useful result.
That demonstrates that I can collaborate with an LLM. It does not establish that the proposed harness can reproduce the interaction. My interpretation and judgment have become hidden parts of the system.
The rule cannot simply be “never intervene,” because a real harness is allowed to enforce rules and manage failures. It can reject an unavailable tool, stop after a defined number of attempts, or ask a user for missing information. Those are behaviors I can implement. The important question is what knowledge I used to decide what should happen.
If I stop after a preset retry limit, I am simulating software. If I stop because my experience tells me the entire approach is misguided, I have introduced a judgment that still needs an owner. It might belong to the model, a separate check, or a human reviewer. Until I decide, I should not count it as solved.
For me, the useful constraint is therefore to make every intervention visible. What triggered it? What information did I use? Could I describe the response precisely enough to implement it, or did it require expertise that the proposed system does not have?
The cricket experiment that did not pass
I also tried this approach with the idea of predicting cricket match scores. I worked with LLMs, supplied the information they requested, and eventually exposed Cricsheet data through Python. Cricsheet provides ball-by-ball cricket match data, so I had moved beyond simply asking a model to produce a prediction from whatever it already knew.
But the interaction kept looping. The model asked for information, I provided it, and the process continued without reaching the result I needed. Even while acting as the enabler, I could not make the approach work. For my decision at that point, it failed the entry test: I did not have enough evidence to justify building a harness around it.
There is a limit to that conclusion. My attempt did not prove that cricket prediction is impossible, or that nobody could build a useful system for it. It showed that the combination of model, tools, instructions, and approach I tried was not giving me a workable path forward.
Looking back, I would make the target more precise before another attempt. Predicting a final score before a match and estimating it halfway through an innings are different tasks. I would also need to define an acceptable error and compare predictions against a baseline on matches whose outcomes were withheld from the prediction process. Getting the model to return a number would not, by itself, establish predictive value.
One possible direction would be a forecasting model with an LLM helping query data or explain results. That is a different proposal to investigate, not a capability I can assume exists because I have connected an LLM to Python.
The entry test still served its purpose. I did not need to prove that every possible approach would fail. I needed a reason to invest in this one, and the experiment had not supplied it.
The student harness has a different success condition
The education project makes the other side of the test clearer. Here, questions are part of the intended behavior. A request for more information could help a student examine their reasoning, or it could leave them stuck in an unproductive exchange. Counting questions would tell me very little about whether the system is doing its job.
Consider an illustrative case: a student writes a loop with an incorrect boundary. A useful response might ask them to trace the final iteration and identify which element the code accesses. If they still cannot see the issue, the next response might provide a smaller example or explain the relevant concept. The challenge is deciding how much support to give and when to give it.
That is the interaction I want to explore before encoding the surrounding workflow. Does the model respond to the student’s actual explanation? Does it recognize a misconception? Can it provide enough information to make progress without automatically completing the exercise? These are questions for the prototype, not results I am claiming for the current project.
There is also a harder evaluation problem. Passing code tests would show something about the program, but would not establish what the student learned. Asking a student to explain their solution or attempt a related problem could provide additional evidence. I would still need to test the educational value with students; a convincing tutoring conversation alone cannot settle it.
In this setting, keeping a person involved is compatible with the product’s purpose. The student is supposed to think and act. The boundary I need to examine is whether I, as the person simulating the harness, am supplying undisclosed teaching decisions that the eventual system cannot reproduce.
Turn interventions into questions before architecture
It is tempting to translate every difficulty directly into a software component: forgotten context means memory, repetition means loop detection, and an incorrect response means a validator. Sometimes those are useful answers. But the intervention tells me where to investigate, not necessarily what to build.
If the model misses an earlier fact, perhaps the context was never supplied clearly. If it repeatedly asks for data, perhaps the task is underspecified or the tool cannot retrieve what it needs. Adding persistent storage or a retry limit may change the behavior without fixing the underlying problem.
For the student project, I would use an intervention log like this. These are possible situations to watch for, rather than findings from a completed evaluation:
| Possible intervention | Question it raises |
|---|---|
| I stop the model from giving the full solution | What assistance is allowed at this stage? |
| I supply an earlier student answer | What learning context must the system retain and retrieve? |
| I replace an unhelpful question | Can the model choose a better question, or is expert review needed? |
| I explain what an assignment expects | Are the objectives and evaluation criteria explicit enough? |
| I end a repetitive exchange | What should trigger a hint, explanation, or escalation? |
| I decide that the student understands | What evidence supports that judgment? |
Some answers may become ordinary application code. Others may require better instructions, better tools, or a deliberate human decision. Discovering that distinction is part of the value of the exercise.
Passing earns a prototype
A successful manual run is encouraging, but it gives me a reason to take the next step. It does not establish that the product will be reliable or economical. I would want to repeat the workflow with different inputs and look at the interventions, unfinished attempts, and failures as well as the successful result. Anthropic’s guidance on evaluating AI agents similarly distinguishes the execution trace from the outcome and emphasizes repeated trials when behavior varies.
For an entry test, I would keep the process small: define one concrete task, decide what useful progress looks like, expose realistic tools, and set a stopping point. Then record what the model accomplishes and what I have to contribute. If I change the instructions or tools after a failure, I should test the revised setup again rather than treating the rescued attempt as evidence that the original setup worked.
The result might justify a small harness. It might suggest a narrower product with human review. It might also reveal that the workflow consists mostly of fixed rules and needs very little model-driven decision-making. Anthropic’s distinction between workflows and agents is useful here: predefined control paths and model-directed actions are different design choices, and added complexity should earn its place.
That is why I prefer to establish the interaction before choosing a framework. Once I understand where decisions belong, I can choose software that helps implement them. A framework can make construction easier; it cannot supply the missing evidence that my proposed workflow is useful.
The rule I actually use
My rule is simple: if I cannot make meaningful progress while manually providing the tools and context the proposed system would have, I pause before building it. I need either a better approach or a clearer reason to continue. If I can make progress, I examine how much of the result depends on judgment I have quietly contributed.
That makes this a personal investment filter. It is intentionally cheaper and less conclusive than evaluating a finished product. The cricket attempt did not give me a reason to proceed with that approach. The student coding harness remains under development, with the important question still open: can the interaction support learning in the way I intend?
If you are considering an AI product, try the exercise on one real workflow. Set your own threshold for useful progress, record your interventions, and let the result influence what you build next. You may discover a promising agent, a simpler automation, or a reason to leave the idea alone for now.
Before building the harness, become the harness—and pay attention to the work you are doing.
LogCTL Dispatch
New essays, experiments and workflows. No daily noise.
Usually 2–4 emails a month.
Comments
Loading comments…