Vector          

Surgical world models need definitive evals.

Vector is Seldinger's internal, open source tool to evaluate our AI work. We standardize how we evaluate physical AI in surgery with pinned tasks and runtimes, task-owned verification, safety gates, and replayable evidence.

01          
Open source

Run Vector on infrastructure you control.

Install the harness, pin the task and verifier, and keep the evidence wherever Vector runs.

GitHub ↗
Hosted

Create an account and start a run.

Use the planner to define the evaluation, attach public or de-identified inputs, and keep each result with its evidence.

Create account →
White glove

We scope and run the evaluation with you.

Fixed-scope protocol, packaged benchmark, reproducible run, failure review, and decision memo.

Schedule a Vector evaluation

White-glove evaluations are scoped with you: protocol, benchmark packaging, reproducible run, evidence review, and a decision memo. Pick a time and we will confirm the workflow and data boundary on the call.

Open Cal.com scheduler
02        

Local, hosted, and Seldinger-run evaluations use the same evidence contract.

  1. 01TaskVersioned instructions, inputs, world, metrics, and gates.
  2. 02AgentDeclared capabilities and an immutable runtime identity.
  3. 03ExecuteLocal, managed, or customer-owned CPU and GPU workers.
  4. 04VerifyA task-owned verifier scores evidence the agent never sees.
  5. 05ReplayThe stored trace must reconstruct the same vector and head.
03                

Inspect the pinned task, model package, scorecard, gate failures, and result head before trusting the claim.