Skip to content
cpbenchy 0.1.0.dev0 is in alpha: until version 1.0, commands, options, the Python API and the result format may still change. Pin the version you use.

Using cpbenchy from your framework

Your framework decides what to run and stores the results; cpbenchy measures the runs, with reliable limits and timings.

Updated View as Markdown

This page is for you if you already have your own way of running experiments, such as a sweep tool, a job database, a lab framework or a pipeline. It already knows which runs exist, which are done and where results go. What you want from cpbenchy is the measuring: running a CPMpy solver on an instance under limits that hold, and reporting the time, memory and answer.

cpbenchy.backend does only that part. You hand it a list of runs, each tagged with your own id, and you get back one result per run, tagged with the same id. What cpbenchy otherwise does for you, such as finding instances, skipping finished runs and storing results, is left to your framework.

Your framework cpbenchy
decides which runs exist, and which still need doing runs each solver in its own process, under its time, CPU time and memory limits
names each run with its own id pins runs to cores and runs them in parallel without disturbing each other
stores the results where it wants measures wall time, CPU time and memory, and reads the solver’s answer
retries, resumes, schedules over machines turns every failure of a run into a result, never an exception

An example

Take a framework that keeps its trials in a SQLite table, and fills in a result for each. Each time you run this script, it measures the trials that have no result yet:

import json
import sqlite3

from cpbenchy import Instance, Limits, RunSpec
from cpbenchy.backend import Submission, run

db = sqlite3.connect("trials.db")   # a table: trials(id, path, solver, seed, result)

# Your framework decides what still needs doing: here, the rows without a result.
todo = db.execute("SELECT id, path, solver, seed FROM trials WHERE result IS NULL").fetchall()

submissions = [
    Submission(
        RunSpec(Instance.from_path(path), solver, Limits(time_s=60, mem_mib=4096), seed=seed),
        key=str(trial_id),
    )
    for trial_id, path, solver, seed in todo
]

def store(item):
    # Called as each run finishes: store it now, so an interrupted call loses nothing.
    db.execute("UPDATE trials SET result = ? WHERE id = ?", (json.dumps(item.result.to_dict()), int(item.key)))
    db.commit()

run(submissions, jobs=4, out="cpbenchy-scratch", on_result=store)

Afterwards, each row holds a result such as {"status": "optimal", "objective": -9, "walltime_s": 1.40, "memory_mib": 55.4, ...}.

The rest of this page goes through the steps: describing a run, what comes back, and what happens when something goes wrong.

Describing a run

A RunSpec says exactly what to measure. It is plain data:

RunSpec(
    instance=Instance.from_path("data/knapsack.opb"),
    solver="ortools",                      # any CPMpy solver name
    limits=Limits(time_s=60, mem_mib=4096, cputime_s=None),
    params={"num_search_workers": 1},      # solver parameters, passed to solve()
    seed=1,                                # None: the solver's default
    cores=1,                               # cores the run gets to itself
    loader=None,                           # None: CPMpy reads the file by its extension
)
  • The instance is a file the worker reads, so its path must exist on the machine that runs the call. Any format CPMpy reads works: XCSP3, OPB, WCNF, DIMACS, MPS and more, compressed or not. Instance.from_path names it after the file. Pass dataset="..." to tell apart instances that share a name, and format="opb" when the extension doesn’t say.
  • Your own formats: write a Loader and pass its reference as loader=, for example "mypkg.loaders:KnapsackJSON".
  • Limits: time_s is wall time and is required. mem_mib and cputime_s are optional. See Measurement and limits for how they are enforced.

Then wrap it in a Submission with your id:

Submission(spec, key="trial-42", metadata={"sweep": "lr-0.1"})

key is any string you use to find the row again. metadata is any dictionary you want back with the result, so on_result doesn’t have to look it up.

Run ids

Each RunSpec has a run_id: a short hash of what it measures. That is the instance, the solver, the parameters, the seed, the cores and the limits. It leaves out your key and metadata. An instance with a dataset counts by its dataset and name, so it keeps its id when its files move. An instance without one counts by its path.

  • Duplicates run once. Submissions with the same run_id are measured once, and each gets the result, under its own key. If ten trials of yours ask for the same measurement, it costs one run.
  • A cache key. The id stays the same between calls, so you can store spec.run_id with each result and look up a measurement you already have before you submit it again.

What comes back

run returns a list with one BackendResult per submission. It calls on_result(item) with each of them as soon as its run finishes, in your own process and thread:

item.key, item.metadata what you submitted
item.result the RunResult: the measurement and the answer
item.log the worker’s output (stdout and stderr), or None; worth keeping for failed runs

The fields of item.result you will most often store:

Field
status optimal, feasible, unsat, unknown, timeout, memout or error
objective the best objective value found, for optimization problems
solved True for an optimal or unsat answer, and for a solution to a satisfaction problem
walltime_s, cputime_s, memory_mib what the run used, measured from outside
parse_s, transform_s, solve_s where the time went, measured in the worker
reliable False when the machine couldn’t enforce limits exactly; see executors
error what went wrong, for status error
extra what plugins recorded

item.result.to_dict() gives all fields as a JSON-ready dictionary, and RunResult.from_dict reads it back. The result record lists every field.

When things go wrong

A failing run is a result, not an exception. A missing file, an unknown solver, a solver that crashes or a model that doesn’t load: each gives a result with status error and a message in error, and the other runs go on. A run that hits its limit gives timeout or memout. So on_result is called once for every submission.

A problem with the whole call raises cpbenchy.errors.UsageError before any run starts. For example, jobs=4 runs of 16 GiB each don’t fit on a 32 GiB machine.

An interrupted call keeps what is done. Ctrl-C stops the runs still going and raises KeyboardInterrupt from run. They get no result, and the runs that had finished already went through on_result. That is why the example stores in on_result rather than from what run returns. Next time, your framework submits what is still missing.

Checking solutions and more

Plugins and observers work here as in any other run. Pass them as plugins=, and their options, or any other command-line option, as args=. What they record ends up in item.result.extra. To check every solution against the model:

run(submissions, plugins=["cpbenchy.observers:CheckSolutions"], on_result=store)
# item.result.extra["check"] == {"valid": True, ...}; a wrong solution makes the status "error"

To let solvers stop at their limit with their best solution, as competitions do, pass args=["--terminate", "--grace", "5"]; see stopping runs at their limit.

Settings

run(
    submissions,
    jobs=1,                    # runs in parallel on this machine
    executor="auto",           # how runs are measured; see Measurement and limits
    plugins=[],                # plugin objects or references
    args=[],                   # any command-line option, including those of plugins
    out=None,                  # cpbenchy's own working directory
    quiet=False,               # True: no live view in the terminal
    on_result=None,
)

run returns when all submissions are done, and on_result runs in between. Keep it quick: while it runs, no new run starts. A database write is fine.

out is where cpbenchy keeps what it needs while it runs: the workers’ logs, and a copy of the results in its own format (results.jsonl). Your framework doesn’t need to read it. Without out, each call makes a new temporary directory and leaves it in place, so pass a directory of your own and delete it when you like. Runs are measured even when out already has them, because deciding what to skip is up to your framework.

Scaling out

One call measures on one machine, using its cores in parallel. Don’t make two calls at the same time on one machine: each places its runs on the cores by itself, so they would share cores and disturb each other’s timings.

To spread runs over several machines, such as the nodes of a cluster, let your framework give each machine its own part of the submissions, and make one call there. The instance files must be readable on each. To run the workers somewhere else entirely, for example in containers, write an executor.

When you don’t need this

If you don’t already have your own bookkeeping, use an Experiment instead. It finds the instances, skips the runs that are done and stores the results, with the same measurements.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close