Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@

- `lemans report -S` sorts by several dash-joined columns (`-S score-credit`); `^column` reverses that column's order (`-S ^score`: low to high, `-S ^model`: Z-A).
- Multistep results record `total_steps`; `lemans report` now shows a `progress` column (steps completed, `2/5`).
- Bring your own agent: `Lemans::Agents.register` adds an agent defined outside lemans, and `lemans run --require FILE` loads the file that registers it.

## [1.5.0] - 2026-10-05

Expand Down
32 changes: 31 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -263,13 +263,43 @@ gpt-5.6-luna ar-archive-book-access 2/2 2m 23s $0.0132 12.5 156905
6 trials: 6 scored, 0 invalid, 6 solved (100%) · $0.0801 · pass@2 3/3 tasks (100%)
```

### Bring your own agent

Any agent that implements the `Lemans::Agent` contract can be measured. It sees its profile (the `agent:` section of `bench.yml`), the task, and the environment, and works through `environment.exec`, `upload`, and `download`. `run` returns an `Agent::Response` with an outcome and usage; `install` runs while the setup network is still open. Define it in a file, register it under a name, and load the file with `--require`:

```ruby
# my_agent.rb
class MyAgent < Lemans::Agent
NAME = "my-agent"

def install(_task, environment)
environment.upload("/path/to/my-agent", "/usr/local/bin/my-agent")
environment.exec!("chmod +x /usr/local/bin/my-agent")
end

def run(task, environment)
finished = environment.exec("cd #{task.environment.workdir} && my-agent --model #{model}", timeout:)
outcome = finished.exit_code == 124 ? :agent_timeout : :completed
Response.new(outcome: Lemans::Result::Outcome.new(outcome), usage: Lemans::Result::Usage.zero)
end
end

Lemans::Agents.register(MyAgent::NAME, MyAgent)
```

```bash
lemans run --bench my-bench --require my_agent.rb --agent my-agent
```

Report what the run cost in `usage` (`cost_usd: nil` when it is unknown), and raise `Lemans::InfrastructureError` or return a `Response` with an `error` when the agent itself failed, so the trial counts as invalid rather than as a zero.

## CLI

| Command | What it does |
| --- | --- |
| `lemans init` | Scaffold a new bench directory: an annotated `bench.yml` and two example tasks |
| `lemans tasks` | List the tasks in a bench (`--tag` to filter) |
| `lemans run` | Run tasks and grade them (`--task`, `--tag`, `--agent`, `--model`, `--max-output-tokens`, `-k`, `-c`, `--resume`) |
| `lemans run` | Run tasks and grade them (`--task`, `--tag`, `--agent`, `--require`, `--model`, `--max-output-tokens`, `-k`, `-c`, `--resume`) |
| `lemans restart <run>...` | Continue failed multistep runs from their last settled step in new runs (`-c`, `--recover` to continue the failed step's session, `--reverify` to grade again, `--allow-scored`, `--backend`, `--max-output-tokens`) |
| `lemans report [RUNS_DIR]` | Summarize `runs/` (or `RUNS_DIR`) as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <columns>` to sort, e.g. `-S score-credit`; numbers high to low, names A-Z, `^column` reverses that column); repeated attempts add pass@k per model × task, fractional grading a `credit` column, multistep tasks a `progress` column (steps completed / task steps) |
| `lemans clobber [RUNS_DIR]` | Delete run results under `runs/` (or `RUNS_DIR`) (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
Expand Down
22 changes: 17 additions & 5 deletions lib/lemans/agents.rb
Original file line number Diff line number Diff line change
Expand Up @@ -11,13 +11,25 @@ module Agents
"miniswen-installed" => "MiniswenInstalled"
}.freeze

def self.build(name, profile:, model: nil) = lookup(name).new(profile: profile, model: model)
@registered = {}

def self.lookup(name)
constant = REGISTRY[name] or
raise ConfigError, "unknown agent #{name.inspect} (known: #{REGISTRY.keys.join(", ")})"
class << self
def register(name, agent_class)
raise ConfigError, "agent #{name.inspect} is built in" if REGISTRY.key?(name)

const_get(constant)
@registered[name] = agent_class
end

def unregister(name) = @registered.delete(name)

def names = REGISTRY.keys + @registered.keys

def build(name, profile:, model: nil) = lookup(name).new(profile: profile, model: model)

def lookup(name)
@registered[name] || (REGISTRY[name] && const_get(REGISTRY[name])) or
raise ConfigError, "unknown agent #{name.inspect} (known: #{names.join(", ")})"
end
end
end
end
6 changes: 5 additions & 1 deletion lib/lemans/cli.rb
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,9 @@ def tasks
option :bench, default: ".", desc: "Directory holding bench.yml"
option :task, desc: "Run task(s) by name", repeatable: true
option :tag, desc: "Run every task carrying this tag(s)", repeatable: true
option :agent, desc: "Override the agent from bench.yml (miniswen, miniswen-installed, oracle, nop)"
option :agent, desc: "Override the agent from bench.yml (miniswen, miniswen-installed, oracle, nop, or one registered by --require)"
option :require, banner: "FILE", repeatable: true,
desc: "Load a Ruby file first, e.g. one that defines an agent and calls Lemans::Agents.register"
option :model, desc: "Override the model(s) from bench.yml", repeatable: true
option :max_output_tokens, type: :numeric, banner: "TOKENS",
desc: "Cap the agent's output per model call (default: the provider's)"
Expand All @@ -58,6 +60,8 @@ def tasks
option :backend, enum: Environments::BACKENDS.keys, desc: "Sandbox backend (default: daytona)"
option :resume, type: :boolean, default: false, desc: "Skip trials that already have a result"
def run_bench
Array(options[:require]).each { require File.expand_path(it) }

# The bundled pricing registry ages faster than the gem: refresh once up
# front, so every trial prices completions against the same revision.
Miniswen.refresh_registry!
Expand Down
43 changes: 43 additions & 0 deletions test/lemans/agents_test.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# frozen_string_literal: true

require "test_helper"

class AgentsTest < Minitest::Test
class TestAgent < Lemans::Agent
NAME = "test-agent"

def run(_task, _environment) = Response.new(outcome: Lemans::Result::Outcome.new(:completed), usage: Lemans::Result::Usage.zero)
end

def teardown
Lemans::Agents.unregister(TestAgent::NAME)
end

def test_register
Lemans::Agents.register(TestAgent::NAME, TestAgent)
config = Lemans::Config::Agent.new(TestAgent::NAME, "some/model")

agent = Lemans::Agents.build(TestAgent::NAME, profile: config, model: "other/model")

assert_instance_of TestAgent, agent
assert_equal "other/model", agent.model
assert_includes Lemans::Agents.names, TestAgent::NAME
assert_instance_of Lemans::Agents::Oracle, Lemans::Agents.build("oracle", profile: config)
end

def test_builtin_names_are_taken
error = assert_raises(Lemans::ConfigError) { Lemans::Agents.register("oracle", TestAgent) }

assert_equal "agent \"oracle\" is built in", error.message
end

def test_unknown_agent
Lemans::Agents.register(TestAgent::NAME, TestAgent)

error = assert_raises(Lemans::ConfigError) do
Lemans::Agents.build("nope", profile: Lemans::Config::Agent.new("nope", "some/model"))
end

assert_equal "unknown agent \"nope\" (known: nop, oracle, miniswen, miniswen-installed, test-agent)", error.message
end
end
33 changes: 33 additions & 0 deletions test/lemans/cli_test.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# frozen_string_literal: true

require "test_helper"
require "lemans/cli"

class CLITest < Minitest::Test
def teardown
Lemans::Agents.unregister("required-agent")
end

def test_require
Dir.mktmpdir do |dir|
file = File.join(dir, "required_agent.rb")
File.write(file, <<~RUBY)
class RequiredAgent < Lemans::Agent
NAME = "required-agent"
end
Lemans::Agents.register(RequiredAgent::NAME, RequiredAgent)
RUBY

_, err = capture_io do
Miniswen.stub(:refresh_registry!, nil) do
assert_raises(SystemExit) do
Lemans::CLI.start([ "run", "--bench", BenchFixture::ROOT.to_s, "--require", file, "--task", "none" ])
end
end
end

assert_includes Lemans::Agents.names, "required-agent"
assert_includes err, "no matching tasks"
end
end
end
Loading