The era of AI agents operating as mere experimental demos is over. As these systems move into production, the industry is suddenly wrestling with a massive question: how do we actually trust them to take action?
This question has birthed an entirely new discipline called AgentOps. Just as DevOps brought structure to software deployment, AgentOps provides the practices, tools, and frameworks required to deploy, monitor, evaluate, and govern autonomous AI agents in production. But of all the pillars of AgentOps, evaluation is proving to be the most critical and the most difficult.
Traditional LLM benchmarks were static, treating AI like a student taking a multiple-choice test. But agents don’t just answer questions; they interact with environments, use tools, and make sequential decisions. For a while, every tech company tried to solve this by building their own proprietary, “walled-garden” testing harnesses. It was fragmented, biased, and impossible to compare apples to apples.
Then I discovered Harbor.
Harbor is an open-source framework designed specifically for evaluating and optimizing agents in sandboxed environments. Rather than forcing developers to build testing infrastructure from scratch, Harbor provides a standardized foundation to spin up agents (like Claude Code or OpenHands) and test their ability to execute commands and complete tasks in secure, isolated containers.
It represents a massive shift toward “shared evaluation infrastructure.” By relying on an open scaffolding like Harbor, developers can focus entirely on designing the actual real-world tests rather than reinventing the sandbox. And this isn’t just a niche tool for indie developers; the biggest players in the cloud ecosystem are already using it as their base layer.
NOTEInterestingly, Amazon Web Services recently used the Harbor framework to create their own open-source project called aws-bench. Instead of just testing agents in a basic virtual environment, aws-bench places AI agents into real, temporary AWS accounts to see if they can fix problems and set up resources without violating security rules or running up massive bills.
Waiting for api.github.com...
Harbor
If we want to evaluate how well AI agents perform in real-world software engineering environments, we need a reliable way to test them. In Harbor, we build these evaluations using tasks, which we can think of as standardized exams for AI agents. In this guide, we will learn how to create our very first Harbor task from scratch by building a practical system administration challenge: generating an SSH key pair. By the end of this tutorial, we’ll know how to
- write clear task instructions,
- set up an isolated test container,
- provide a reference solution to prove solvability, and
- build an automated grading rubric.
Step 1: Create Task
A task defines one or more instructions, a sandbox environment, and a verifier. Tasks are used to evaluate agents and models and are implemented as directories in the Harbor task format.
Think of Harbor task as standardized exam for AI agents that includes:
- the question: The prompt or problem statement given to the AI
- the workbench: an isolated container with necessary software installed as a clean workspace provided for the task
Run the following command to create a new task directory with the required files:
harbor task init ssh-key-pairWe will evaluate/benchmark an agent or model using task like this.
Step 2: Writing Task Instructions
Open the instruction.md file in the task directory and add the task description:
# SSH Key Pair Generation
Generate an SSH key pair in the files `~/.ssh/id_rsa` and `~/.ssh/id_rsa.pub`.
Don't make them password protected.Think of instruction.md as the “exam question”. In this example, our agent will perform a task defined by this
instruction.md file and generate some verifiable results as the agent’s “answer sheet to the exam”.
Step 3: Configuring Task Metadata
Open the task.toml file to configure task metadata, resource limits, and timeouts.
NOTEMost fields in
task.tomlare optional, allowing Harbor to defer configuration to our environment files unless specific overrides or metadata are needed.
version = "1.0"
[metadata]author_name = "Your Name"author_email = "your.email@example.com"difficulty_explanation = "Simple SSH key generation command"category = "system-administration"tags = ["ssh", "cryptography", "linux"]
[verifier]timeout_sec = 120.0
[agent]timeout_sec = 120.0
[environment]build_timeout_sec = 600.0Step 4: Creating Task Container In Which Agent Works
Now that we have exam sheet, let’s create a test room for our agent. A container
Dockerfile defines such an environment an agent will interact with through the terminal. Open the Dockerfile in the
environment/ directory that was generated and add any dependencies our task needs:
FROM ubuntu:24.04
# Create working directoryWORKDIR /app
# Install openssh-client for the taskRUN apt-get update && apt-get install -y openssh-client && rm -rf /var/lib/apt/lists/*Harbor spins up the container using this Dockerfile and the AI agent is injected inside it. The agent will interact
with this container environment via a simulated terminal - running shell commands, installing tools, and editing files
right inside the container
TIPThe AI agent under test runs inside the container.
Step 5: Creating Solution Script
Proof of Solvability (The “Oracle” Agent)Before benchmarking real AI models, we need to verify that our task isn’t broken (e.g., checking that required dependencies like
openssh-clientaren’t missing from the container).To do that, Harbor runs an Oracle Agent, which simply executes some script called solution script inside the test container followed by the test script to verify the outcome.
If such script passes, we can prove that our environment (“test room”) is good for the agent to solve the task (“finish the exam”)
We implement this reference solution in ssh-key-pair/solution/solve.sh below:
#!/bin/bash
ssh-keygen -t rsa -f ~/.ssh/id_rsa -N ""We also want to make sure the script is executable:
chmod +x ssh-key-pair/solution/solve.shNow we can run the following command to verify our task is solvable by the solution script:
cd ../ # navigate to the directory that contains the "ssh-key-pair" task folderharbor run -p ssh-key-pair -a oracleIf successful, we should see output indicating the task was completed and the reward was 1.
Step 6: Creating Test Script
Now that the tested agent has “finished the exam”, we need to grade the agent’s answer, i.e. whether the agent has successfully completed the task. We do this via test scripts, which consist of 2 parts:
- A language-agnostic test harness that “grades agent’s answer”
- An executor that runs the test harness and must be written in bash named
test.sh
We first write the actual tests, for example in Python for simplicity:
import osfrom pathlib import Path
def test_key_files_exist() -> None: """Test that both private and public key files exist.""" private_key = Path.home() / ".ssh" / "id_rsa" public_key = Path.home() / ".ssh" / "id_rsa.pub"
assert private_key.exists(), "Private key file does not exist" assert public_key.exists(), "Public key file does not exist"
def test_key_file_permissions() -> None: """Test that the key files have correct permissions.""" private_key = Path.home() / ".ssh" / "id_rsa" public_key = Path.home() / ".ssh" / "id_rsa.pub"
private_perms = oct(os.stat(private_key).st_mode)[-3:] public_perms = oct(os.stat(public_key).st_mode)[-3:]
assert private_perms == "600", ( f"Private key has incorrect permissions: {private_perms}" ) assert public_perms == "644", ( f"Public key has incorrect permissions: {public_perms}" )
def test_key_format() -> None: """Test that the public key has the correct RSA format.""" public_key = Path.home() / ".ssh" / "id_rsa.pub"
with open(public_key, 'r') as f: content = f.read()
assert content.startswith("ssh-rsa "), "Public key does not start with 'ssh-rsa'" assert len(content.split()) >= 2, "Public key format is invalid"Next, create the execution wrapper that must be written in Bash and named test.sh:
IMPORTANTThe script must write a numerical reward (
1for pass,0for fail) to/logs/verifier/reward.txt. This is how Harbor communicates whether the agent succeeded.
#!/bin/bash
apt-get updateapt-get install -y curl
curl -LsSf https://astral.sh/uv/0.9.5/install.sh | sh
source $HOME/.local/bin/env
# Run pytest testsuvx \ --python 3.12 \ --with pytest==8.4.1 \ pytest /tests/test_outputs.py
# Check exit code and write rewardif [ $? -eq 0 ]; then echo 1 > /logs/verifier/reward.txtelse echo 0 > /logs/verifier/reward.txtfiSolution Script v.s. Test Script
It is very easy to get lost on why we need to have both solution script and test script.
Although comparing an agent’s output directly against solve.sh sounds logical on the surface, in complex environments
solve.sh and test.sh serve 2 completely different purposes: Execution vs. Validation.
Having both is essential because solve.sh is one way to do the work while test.sh checks if the work has done
An AI agent will almost never write the exact same script or run the exact same steps as the solve.sh, which contains
the reference commands required to solve the task.
test.sh, on the other hand, contains the assertion logic that evaluates the final state of the system regardless of
how the agent accomplished it. In software engineering, there are usually many ways to achieve the same result:
- One agent might run
ssh-keygen -t rsa -N "" -f ~/.ssh/id_rsa. - Another agent might write a small Python script using
paramikoto generate the key. - A third agent might run
cat << 'EOF' > generate.sh ...and execute it.
solve.sh represents just one valid path to the solution. test.sh evaluates the final state of the system,
allowing the agent full freedom in how it solves the task.
If Harbor only compared the agent’s actions to solve.sh, any agent that used a slightly different syntax, ran
intermediate debugging commands, or wrote a custom script would fail—even if its final result was 100% correct.
| Component | Exam Analogy | Role |
|---|---|---|
solve.sh | The Teacher’s Sample Solution | Shows one worked-out way to solve the problem to prove it is doable. |
test.sh | The Automated Grader / Rubric | Assesses whether the final outcome meets all functional criteria, no matter how the student got there. |