Skip to content
cpbenchy 0.1.0.dev0 is in alpha: until version 1.0, commands, options, the Python API and the result format may still change. Pin the version you use.

PAR-k scores

Ranks solvers the way competitions do: solved fast is good, unsolved is penalised.

Updated View as Markdown
OptionBuilt incomes with cpbenchy
cpbenchy run instances/ -s ortools -t 60 --par 2
Version
cpbenchy 0.1.0.dev0
Last updated
6 Oct 2026 · 1 commit
Authors
ThomSerg
Tags
scoringcompetitionpar2

What it does

Ranks solvers with one number, as many solver competitions do, such as the SAT competitions with PAR-2. PAR-k (penalised average runtime) scores each run, and a solver’s score is the total over its runs. Lower is better.

  • A solved run scores its time. It is solved when it ends optimal or unsat, or feasible on a problem without an objective.
  • Any other run scores k times its time limit: a timeout, a memout, an error, or a solution not proven optimal.

So a solver gains more by solving one more instance than by solving the others a bit faster, and k says by how much.

Runs with a CPU time limit, as under the PB rules, are scored by CPU time, against that limit. Other runs are scored by wall time.

Use it

--par adds a column to the summary:

┏━━━━━━━━━┳━━━━━━┳━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┓
┃ solver  ┃ runs ┃ solved ┃ optimal ┃ timeout ┃ time solved ┃ PAR-2 ┃
┡━━━━━━━━━╇━━━━━━╇━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━┩
│ exact   │   10 │      7 │       7 │       3 │       52.3s │ 412.3 │
│ ortools │   10 │      6 │       6 │       4 │       58.9s │ 538.9 │
└─────────┴──────┴────────┴─────────┴─────────┴─────────────┴───────┘

Scores are computed from the stored results, so you can score an experiment after the fact, with any k. Compare totals over the same instances only: a solver with fewer runs has a lower total.

Options

--par K the penalty for an unsolved run, as a multiple of its time limit. Also par = 2 in cpbenchy.toml or in rules
par(result, factor=2, time=None) one run’s score; time="walltime" or "cputime" to choose which time counts
par_totals(results, factor=2, time=None, by=("solver",)) the total per group, as a dictionary keyed by the values of by

Rules that set par stay followed when you choose another k: it changes how results are reported, not what is measured.

Implementation

The par in src/cpbenchy/scoring.py, lines 23–39 of 53, as of this version of the docs.

src/cpbenchy/scoring.pypython
def par(result: RunResult, factor: float = 2, time: Time | None = None) -> float:
    """The PAR-`factor` score of one run: its time if solved, else `factor` times its time limit.

    `time` is "walltime" or "cputime". Without it, a run with a CPU time limit is scored by CPU time
    against that limit, as competitions with CPU time limits do, and any other run by wall time.
    """
    if time is None:
        time = "cputime" if result.cputime_limit_s is not None else "walltime"
    if time == "cputime":
        used, limit = result.cputime_s, result.cputime_limit_s or result.time_limit_s
    elif time == "walltime":
        used, limit = result.walltime_s, result.time_limit_s
    else:
        raise ValueError(f"time is 'walltime' or 'cputime', not {time!r}")
    if result.solved and used is not None:
        return used
    return factor * limit
Navigation

Type to search…

↑↓ navigate↵ selectEsc close