Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
99 commits
Select commit Hold shift + click to select a range
e6f377b
Add harness-engineering-bench (benchmarks at top level)
varunursekar Jul 23, 2026
6c514f4
gaia: baseline_floor off (feedback #1)
varunursekar Jul 23, 2026
2a8b2cf
gaia: re-enable harness isolation (harness_user: harness)
varunursekar Jul 23, 2026
a70c572
Bump harbor to 0.20.0 across benchmark configs and regenerate locks
varunursekar Jul 23, 2026
801ca6a
Benchmarks: adopt .evals context naming and the evals CLI
varunursekar Jul 24, 2026
42e6cac
Benchmarks: run inner eval sandboxes in a dedicated Modal app
varunursekar Jul 24, 2026
4ad1855
swe-atlas-qna: submit-mode selection, matching gaia
varunursekar Jul 24, 2026
c7bce9a
Benchmarks: enable W&B reporting on officeqa, swe-atlas-qna, and tau3
varunursekar Jul 24, 2026
2543a06
swe-atlas-qna: route the in-container rubric judge to the upstream
varunursekar Jul 24, 2026
3387e12
officeqa: merge duplicate extra_harbor_args keys
varunursekar Jul 24, 2026
43c7591
officeqa: submit-mode selection, matching gaia
varunursekar Jul 24, 2026
45f69bd
Benchmarks: bound orphaned infrastructure from killed runs
varunursekar Jul 24, 2026
7f32aa2
Benchmarks: normalize optimization task configurations
varunursekar Jul 24, 2026
82fcbd7
Benchmarks: document the normalized configuration
varunursekar Jul 24, 2026
6cf6d06
swe-atlas-qna: keep Modal sandboxes alive under the images' ENTRYPOINT
varunursekar Jul 24, 2026
a91377b
Add post-hoc per-trial token aggregator
varunursekar Jul 24, 2026
20650c1
Docs: note infra retry/exclusion is trusted-only in benchmark config
varunursekar Jul 24, 2026
60d98f7
Docs: add target model as a per-benchmark field in CONFIGURATION.md
varunursekar Jul 25, 2026
6d292b5
swe-atlas: port baseline agent to Chat Completions; best-effort answe…
varunursekar Jul 25, 2026
c0df0e4
browsecomp-plus: reduce partitions to 33/66/66 (in line with the suite)
varunursekar Jul 25, 2026
b4a6dcf
officeqa + tau3: port baseline agents to Chat Completions
varunursekar Jul 25, 2026
3ef4fb8
officeqa: gate parallel_tool_calls to non-fireworks models (as reason…
varunursekar Jul 25, 2026
2d04789
officeqa: best-effort answer at turn cap instead of raising
varunursekar Jul 25, 2026
9812985
Add BrowseComp-Plus benchmark (deep-research over a fixed BM25 corpus)
varunursekar Jul 25, 2026
5359949
browsecomp-plus: placeholder OPENAI_API_KEY so pyserini's BM25 import…
varunursekar Jul 25, 2026
51013e1
browsecomp-plus: pass judge credentials via [verifier.env]
varunursekar Jul 25, 2026
756fc6b
Set deepseek-v4-flash as default target model (gpt-oss-120b on swe-at…
varunursekar Jul 25, 2026
653648d
Housekeeping: drop stale gitignore note, fix README candidates table,…
varunursekar Jul 25, 2026
cde12b1
Add in-place 429 retry (max_retries=8) to the four baseline agents
varunursekar Jul 25, 2026
beeaa01
nit: style
varunursekar Jul 25, 2026
d078f75
Pin held-out baselines (K=3) into baseline_reward; disable score_base…
varunursekar Jul 25, 2026
2905da4
Drop the ALE-Bench benchmark
varunursekar Jul 26, 2026
55bd0f9
Assert benchmark-config invariants from the configs themselves
varunursekar Jul 26, 2026
7434a13
Act on the PR review: benchmark and baseline-agent fixes
varunursekar Jul 26, 2026
9e5e198
Skip the benchmark-config tests only for tasks nobody vendored
varunursekar Jul 26, 2026
51c1e4d
Spell the producer default with its provider prefix
varunursekar Jul 26, 2026
a089898
Give every baseline agent's tool dispatch the same failure envelope
varunursekar Jul 26, 2026
4e90dac
gaia agent: don't crash on reason/search-only turns; force-final at cap
varunursekar Jul 26, 2026
cce53db
Pin gaia held-out baseline (0.5736) with score_baseline: false
varunursekar Jul 26, 2026
5590f8f
officeqa: average the noisy held-out finalize over 3, single W&B project
varunursekar Jul 27, 2026
aae5bdb
Add a run-benchmark skill: how to launch and health-check a run
varunursekar Jul 27, 2026
2504532
Re-time case/verifier timeouts from baseline waits; fix officeqa opti…
varunursekar Jul 27, 2026
368f435
Capture per-trial token/latency metrics; enable trusted attribution o…
varunursekar Jul 27, 2026
6259be1
Report per-trial token and latency distributions in the aggregator
varunursekar Jul 27, 2026
635ff26
Give finalization its own gateway budget, sized so cases bind before …
varunursekar Jul 27, 2026
f23fa63
Size the verifier's clocks to be unreachable, and raise case concurre…
varunursekar Jul 27, 2026
eb8774c
Take per-case timeouts from each dataset instead of choosing our own
varunursekar Jul 27, 2026
3712a52
Correct stale execution defaults in the configuration doc
varunursekar Jul 27, 2026
052c63d
Average the held-out score over 3 on every benchmark, not just officeqa
varunursekar Jul 27, 2026
00d26a2
Add a salvage path to re-score a candidate after a verifier failure
varunursekar Jul 27, 2026
17e8ef8
Document the opencode launch recipe, which silently bypasses the gate…
varunursekar Jul 27, 2026
c13b45f
Add a credential-file template and document how to run benchmarks con…
varunursekar Jul 27, 2026
8112bd4
Measure the real concurrency ceiling instead of reasoning about it
varunursekar Jul 27, 2026
1d41545
Correct the opencode recipe: anthropic/ now works, openai/ crashes on…
varunursekar Jul 27, 2026
59e2a9b
Let uv-installed optimizer harnesses install as the unprivileged agen…
varunursekar Jul 27, 2026
460dbcc
Point the remaining three benchmarks at the shared W&B project
varunursekar Jul 28, 2026
ab8717b
Tell the runbook to copy the session out before killing a run
varunursekar Jul 28, 2026
ec07dde
Correct the concurrency ceiling: it is per LiteLLM key, and keys scale
varunursekar Jul 28, 2026
2603a07
Add a checker for the per-key rate limits that bound concurrency
varunursekar Jul 28, 2026
140f684
Note that raised key limits take ~15 min to propagate
varunursekar Jul 28, 2026
f7d140b
feat: scaffold SWE-bench-Pro baseline for the vero optimizer
shehabyasser-scale Jul 24, 2026
bd68ebf
fix: spell swe-bench-pro's models the way the gateway matches them
shehabyasser-scale Jul 26, 2026
227da77
feat: preflight configured models before launching an optimizer trial
shehabyasser-scale Jul 25, 2026
6667e3d
fix: do not read a missing route as a missing model in preflight
shehabyasser-scale Jul 26, 2026
dec1ecd
fix: classify Modal sandbox / tests-dir loss as transient_infra
shehabyasser-scale Jul 25, 2026
d99d6f4
fix: a missing upstream deployment is not a task failure
shehabyasser-scale Jul 25, 2026
e14eaff
feat: pin swe-bench-pro to swebenchpro@1.0 and generate the real split
shehabyasser-scale Jul 27, 2026
3d58ea7
fix: repair five swe-bench-pro seed-agent defects that corrupt the re…
shehabyasser-scale Jul 27, 2026
a4e956a
test: cover swe-bench-pro in the benchmark-config invariants
shehabyasser-scale Jul 27, 2026
1aec11f
fix: only send reasoning.effort to models that support it
shehabyasser-scale Jul 25, 2026
5bb029a
fix: gate reasoning.effort on capability in the three ported agents
shehabyasser-scale Jul 26, 2026
c156c59
fix: pick the token-limit parameter by model capability too
shehabyasser-scale Jul 26, 2026
a3e0455
fix: route optimizer-trial harbor flags so tau3 survives teardown
shehabyasser-scale Jul 25, 2026
aa6a063
fix: stop claiming a tau3 fix no experiment supports
shehabyasser-scale Jul 25, 2026
5f63149
test: teach the run-command config stubs about optimizer_harbor_args
shehabyasser-scale Jul 26, 2026
934cc4f
fix: repair a scheme-less WANDB_BASE_URL and stop losing sink failures
shehabyasser-scale Jul 25, 2026
2aff9f4
Merge pull request #54 from scaleapi/fix/judge-reasoning-effort
shehabyasser-scale Jul 28, 2026
c8b6871
Merge pull request #55 from scaleapi/fix/tau3-trial-resilience
shehabyasser-scale Jul 28, 2026
16fbc2a
Merge pull request #57 from scaleapi/fix/wandb-base-url-normalization
shehabyasser-scale Jul 28, 2026
d93c696
Merge pull request #50 from scaleapi/feat/swe-bench-pro-baseline-scaf…
shehabyasser-scale Jul 28, 2026
5727a44
Let the rescore script measure a benchmark's seed harness, not just a…
varunursekar Jul 28, 2026
44a852e
Re-pin all five held-out baselines from fresh 3-round measurements
varunursekar Jul 28, 2026
ed55bc7
Merge pr2-add-vero into pr3-harness-bench to end the vero/src divergence
varunursekar Jul 28, 2026
7fed76c
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
d34bf7b
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
315eed6
Recover producer tokens from runs metered before the cache fix
varunursekar Jul 28, 2026
dde22c7
Stop check_keys pointing concurrency planning at core count
varunursekar Jul 28, 2026
68ab9b2
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
3d505c9
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
2d7ab39
Document harness routing and drop the internal W&B host
varunursekar Jul 28, 2026
7ea8772
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
838374a
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
e77e36f
Reserve both spellings of the flags vero owns
varunursekar Jul 28, 2026
78267d6
Keep browsecomp-plus target inference on the gateway
varunursekar Jul 28, 2026
fef429a
Stop the docs claiming guarantees the code does not make
varunursekar Jul 28, 2026
00eb881
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
8336a10
Record that the pinned baseline now reaches the run
varunursekar Jul 28, 2026
19e03b4
Merge branch 'pr2-add-vero' into pr3-harness-bench
varunursekar Jul 28, 2026
3bae4a8
Cover the unprovisioned-model category and fix a wrong annotation
varunursekar Jul 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 0 additions & 3 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,3 @@ wheels/
secrets.env
*.secrets.env
!*.example

# Local benchmark scoping notes (paper planning; not tracked)
harness-engineering-bench/benchmark-scoping.md
3 changes: 3 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
[submodule "browsecomp-plus-upstream"]
path = harness-engineering-bench/browsecomp-plus/upstream
url = https://github.com/texttron/BrowseComp-Plus.git
437 changes: 437 additions & 0 deletions harness-engineering-bench/CONFIGURATION.md

Large diffs are not rendered by default.

50 changes: 50 additions & 0 deletions harness-engineering-bench/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
# Harness engineering benchmarks

`harness-engineering-bench` contains end-to-end benchmarks for automatically
improving the harness of an agent or the code of a component used to build
agents. Each leaf directory pairs one editable target program with one immutable
Harbor dataset and compiles them into an outer Harbor optimization task.

The benchmark definitions intentionally keep three boundaries visible:

- `target/` is the program the optimization agent may edit.
- `partitions/` pins the cases and the development/validation/test split.
- `build.yaml` is trusted configuration: model, evaluator, access policy,
budgets, and final scoring.

In each benchmark, the complete development tasks and attachments are mounted
read-only for the optimization agent. Development evaluations expose per-case
results and complete Harbor trial records, including exact failures and
target-agent logs. Validation remains aggregate-only, and
test is reachable only by the trusted final verifier.

The paper-era benchmark stack remains available on the `paper/v1` branch and
the `paper-v1` tag. New benchmarks should use this Harbor-native layout.

## Benchmarks

Promoted benchmarks live at the top level. Task sets still under review live in
`candidates/`; we work through the list in the paper's `benchmark-scoping.md` and
promote a task set to the top level once it is ready.

### Promoted

| Benchmark | Editable target | Dataset | Split |
| --- | --- | --- | --- |
| [GAIA baseline](gaia/baseline/) | Tool-using Responses API agent | Harbor `gaia/gaia` | 20% / 40% / 40% |
| [OfficeQA baseline](officeqa/baseline/) | Grounded document-QA agent | Treasury Bulletin corpus | 20% / 40% / 40% |
| [SWE-Atlas-QnA baseline](swe-atlas-qna/baseline/) | Codebase investigation agent | Harbor `scale-ai/swe-atlas-qna` | 20% / 40% / 40% |
| [tau3 baseline](tau3/baseline/) | MCP customer-service agent | Harbor `sierra-research/tau3-bench` | 20% / 40% / 40% |
| [BrowseComp-Plus baseline](browsecomp-plus/baseline/) | Fixed-corpus deep-research agent | Pinned local Harbor tasks | 20% / 40% / 40% |

**`swe-bench-pro/` is at the top level but is not promoted, and its numbers are
not comparable to the five above.** It predates the normalization pass those five
went through and still differs on most of it: case budgets are 1x the partition
size rather than 4x, the held-out target has no `n_attempts: 3` / `mean`
override so it is scored once, there is no pinned `baseline_reward` (and
`score_baseline: true` adds a second full held-out pass), the agent clock runs at
0.6x the declared case timeout, gateway request and token caps are 20-30x tighter
than the sizing convention, there is no `agent_env` block, and telemetry goes to
its own W&B project. Treat it as a work in progress: launching it will produce a
number, but not a measurement of the same quantity. `CONFIGURATION.md` documents
the conventions it is missing.
6 changes: 6 additions & 0 deletions harness-engineering-bench/browsecomp-plus/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# Generated from the obfuscated upstream dataset. These files contain decrypted
# benchmark questions and answers and must not be committed.
tasks/
baseline/target/.venv/
**/__pycache__/
**/*.pyc
53 changes: 53 additions & 0 deletions harness-engineering-bench/browsecomp-plus/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# BrowseComp-Plus

This benchmark turns all 830 queries from
[BrowseComp-Plus](https://github.com/texttron/BrowseComp-Plus) into local Harbor
tasks. It evaluates a deep-research agent against the benchmark's fixed corpus
and canonical BM25 index rather than the live web.

## Pinned sources

- Upstream repository submodule: `046949032b0328319cc9a02663a759ec601d9402`
- Query dataset `Tevatron/browsecomp-plus`:
`144cff8e35b5eaef7e526346aa60774a9deb941f`
- BM25 index `Tevatron/browsecomp-plus-indexes`:
`b3f37f70c33829eb09d04784a54277a31871fd63`

The submodule is the Git pointer requested for the integration. The generator
refuses to run if it is checked out at another commit. Hugging Face revisions
are full immutable commit ids as well.

## Generate Harbor tasks

Initialize the submodule and run the builder from this directory:

```bash
git submodule update --init harness-engineering-bench/browsecomp-plus/upstream

cd harness-engineering-bench/browsecomp-plus
uv run --no-project --python 3.12 --with datasets==4.0.0 -- \
python scripts/build_tasks.py
```

The builder downloads and decrypts the pinned query dataset using the pinned
upstream implementation, then writes one complete Harbor task per query under
the ignored `tasks/` directory. It also regenerates the committed deterministic
166/332/332 development, validation, and test split. Use `--force` to replace
an existing generated tree or `--check` to verify it byte-for-byte.

The first Harbor image build downloads the pinned BM25 index (about 2.2 GB).
Every task has an identical environment, so subsequent tasks reuse that image.

## Scoring and trust boundary

Answers use BrowseComp-Plus's required Explanation / Exact Answer / Confidence
format. The verifier follows the official upstream OpenAI evaluator and its
default `gpt-4.1` judge. The upstream primary leaderboard instead runs the same
judge prompt with Qwen3-32B; this Harbor integration deliberately uses the
repository's supported API evaluator so tasks do not require a local GPU judge.

The judge runs inside the task environment and therefore receives the real
upstream credential. As with the SWE-Atlas-QnA rubric judge and tau3 task-owned
LLM services, this disables uid isolation and assumes a non-adversarial
optimizer. The editable baseline itself exposes no live-web or shell tool and
uses only the fixed local index.
3 changes: 3 additions & 0 deletions harness-engineering-bench/browsecomp-plus/baseline/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
compiled/
target/.venv/
target/.pytest_cache/
20 changes: 20 additions & 0 deletions harness-engineering-bench/browsecomp-plus/baseline/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# BrowseComp-Plus baseline

This editable target is a Responses API deep-research agent with three tools:
search the pinned BM25 index, open a document, and submit a formatted response.
The optimization agent may change its prompts, control flow, tool use, or
dependencies, but not the dataset, index, split, evaluated model, or verifier.

Build the generated tasks first as described in the parent
[`README.md`](../README.md), then compile from the repository root:

```bash
cd vero
VERO_SKIP_SECRET_CHECK=1 uv run vero harbor build \
--config ../harness-engineering-bench/browsecomp-plus/baseline/build.yaml \
--output ../harness-engineering-bench/browsecomp-plus/baseline/compiled
```

For a real run, copy `secrets.env.example` to the ignored `secrets.env`, fill it
in, and use `vero harbor run` in the same way as the other harness-engineering
benchmarks.
148 changes: 148 additions & 0 deletions harness-engineering-bench/browsecomp-plus/baseline/build.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,148 @@
name: vero/optimize-browsecomp-plus-baseline
description: >-
Improve a deep-research agent on BrowseComp-Plus using its fixed corpus,
canonical BM25 index, and semantic answer evaluator.
agent_repo: target
task_source: ../tasks
task_manifest: ../partitions/manifest.json
agent_import_path: browsecomp_plus_agent.agent:BrowseCompPlusAgent
harbor_requirement: harbor[modal]==0.20.0

partition_files:
development: ../partitions/development.json
validation: ../partitions/validation.json
test: ../partitions/test.json

# Four full passes over each agent-visible partition.
agent_access:
- partition: development
disclosure: full
expose_case_resources: true
total_runs: 100
total_cases: 132
- partition: validation
disclosure: aggregate
expose_case_resources: false
min_aggregate_cases: 5
total_runs: 100
total_cases: 264

selection_partition: validation
targets:
- partition: test
reward_key: reward
baseline_reward: 0.4619 # re-pinned 0.424 / 0.470 / 0.492 (sd 0.0283); was 0.4495. See runs/BASELINES.md
failure_value: 0.0
max_attempts: 1
# The held-out eval is noisy: score the selected candidate 3x per case and
# average, so the final reward is comparable to the pinned baseline, which was
# itself pooled over 3 rounds. Without this the candidate carries ~sqrt(3) more
# standard error than the floor it is judged against.
# Per-target override - search/validation keep the global n_attempts (1).
n_attempts: 3
aggregate_attempts: mean

evaluation_set_name: browsecomp-plus
objective:
selector:
metric: score
direction: maximize
reward_mode: submit
baseline_floor: false
score_baseline: false
rescore_top_k: 3
rescore_attempts: 1

model: fireworks_ai/deepseek-v4-flash
environment_name: ${inner_env:-modal}
extra_harbor_args: ["--ek", "app_name=harness-engineering-bench", "--ek", "sandbox_idle_timeout_secs=3600"]
harbor_python_version: "3.12"
n_attempts: 1
max_retries: 1
infrastructure_max_attempts: 3
infrastructure_retry_delay_seconds: 5
aggregate_attempts: best
feedback_transcripts: true
feedback_max_bytes: 16000
expose_attempt_detail: false
# Unreachable: worst case is ceil(198/24) x 3600 = 32400s, every
# finalize trial (66 held-out x n_attempts=3) hitting its own cap. Assumes
# max_concurrency=24; recompute if that drops.
timeout_seconds: 39600
# Exactly the dataset's declared [agent] timeout_sec, so vero's derived
# --agent-timeout-multiplier is 1.0 and the target agent gets precisely the
# clock the benchmark intends. Harbor times agent setup, environment build and
# verification on separate clocks with separate multipliers, so none of them
# eat into this budget and no buffer is warranted.
case_timeout_seconds: 3600
task_agent_timeout_seconds: 3600
max_concurrency: 24 # 8 -> 24; see officeqa for the measured headroom argument
error_rate_threshold: 0.1
# Unreachable: worst-case finalize (32400) + worst-case rescore_top_k=3
# validation rescore (32400). A verifier timeout loses the score outright.
verifier_timeout_seconds: 75600
secrets:
- MODAL_TOKEN_ID
- MODAL_TOKEN_SECRET
- WANDB_API_KEY
- WANDB_BASE_URL

wandb:
project: harness-engineering-bench # one project for the whole suite
group: browsecomp-plus # keeps the benchmark distinguishable in the shared project
name: ${wandb_run:-browsecomp-plus} # per-launch label, e.g. --param wandb_run=browsecomp-plus__claude-sonnet-5
tags: [browsecomp-plus]
log_traces: true

inference_gateway:
upstream_api_key_env: OPENAI_API_KEY
upstream_base_url_env: OPENAI_BASE_URL
producer:
allowed_models: ["${optimizer_model:-openai/gpt-5.4}"]
max_concurrency: 8
# See officeqa/baseline/build.yaml for the sizing rationale: the case budget is
# the spend control, so a token cap only needs to stop a runaway.
evaluation:
allowed_models: [fireworks_ai/deepseek-v4-flash]
max_requests: 200000
max_tokens: 2000000000 # 396 agent case-runs (132 dev + 264 validation)
max_concurrency: 64
# Reserved so a search-phase overspend can never starve held-out scoring.
finalization:
allowed_models: [fireworks_ai/deepseek-v4-flash]
max_requests: 200000
max_tokens: 2000000000 # 66 test cases x3 attempts + rescore headroom
max_concurrency: 64

instruct_multifidelity: true
instruct_exhaust_budget: true

# The official upstream OpenAI evaluator runs inside the task container. It
# uses the raw upstream credential while the editable target agent continues to
# receive only the evaluation-scoped gateway credential.
task_services_use_upstream: true
harness_user: null
# Optimizer-agent env (forwarded to the harbor claude-code agent as --ae KEY=VALUE).
# Claude Code's Bash tool caps a single call at BASH_MAX_TIMEOUT_MS (default
# 600000=10min), well under one inner eval, which pushed the officeqa optimizer
# into --detach + background-poll + end-turn -- and a headless --print run is
# never re-woken, so the search died there. Raise the cap so a whole eval fits in
# one blocking call. The background-task vars are defence in depth only: they gate
# *automatic* backgrounding and do NOT remove the Bash tool's run_in_background
# parameter, which the model can still choose. The instruction forbids that.
agent_env:
# Above this benchmark's widest single eval: a full validation pass is
# ceil(66/24) x 3600 = 10800s worst case.
BASH_MAX_TIMEOUT_MS: "14400000"
BASH_DEFAULT_TIMEOUT_MS: "14400000" # same as max: an un-timed eval must still block
ENABLE_BACKGROUND_TASKS: "0"
FORCE_AUTO_BACKGROUND_TASKS: "0"
# Harnesses installed with `uv tool install` (mini-swe-agent, swe-agent)
# default to symlinking their entry point into /usr/local/bin, which the
# unprivileged optimizer user cannot write: "Failed to install executable
# ... Permission denied". npm/nvm-based harnesses (claude-code, opencode)
# are unaffected, so this only bites when the harness changes.
UV_TOOL_BIN_DIR: "/home/agent/.local/bin"

task_environment:
BROWSECOMP_JUDGE_MODEL: gpt-4.1
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# Optimizer and evaluated-agent inference. OPENAI_BASE_URL may point to an
# OpenAI-compatible upstream; use https://api.openai.com/v1 for OpenAI.
OPENAI_API_KEY=sk-your-upstream-inference-key
OPENAI_BASE_URL=https://api.openai.com/v1

MODAL_TOKEN_ID=your-modal-token-id
MODAL_TOKEN_SECRET=your-modal-token-secret

WANDB_API_KEY=your-wandb-api-key
WANDB_BASE_URL=https://api.wandb.ai
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
[project]
name = "vero-browsecomp-plus-agent"
version = "0.1.0"
description = "Editable Harbor-native BrowseComp-Plus baseline for VeRO"
requires-python = ">=3.12"
dependencies = [
"harbor==0.20.0",
"openai==2.46.0",
]

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

[tool.hatch.build.targets.wheel]
packages = ["src/browsecomp_plus_agent"]

[dependency-groups]
dev = [
"pytest>=9.0.2",
"pytest-asyncio>=1.3.0",
]

[tool.pytest.ini_options]
testpaths = ["tests"]
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
"""Harbor-native BrowseComp-Plus target agent."""

from browsecomp_plus_agent.agent import BrowseCompPlusAgent

__all__ = ["BrowseCompPlusAgent"]
Loading