Model Trials
A small experiment, kept honest
Same brief. Different paths.
Collect plans, choose what to build, then compare the work. You run the models; this workspace keeps the experiment together.
Workspace files
Synthetic example · fictional model identifiers and invented results. Edit freely; this is not evidence about real models.
What should good look like?
One task per workspace. Put tools, time budget, starting files and constraints in the brief so each run starts on the same footing.
1–5 criteria, up to 500 characters each. Rate each 1–5: 1 misses, 3 partly meets, 5 fully meets. Blank means unjudged.
Up to 10 per role, 200 characters each. Include the configuration you actually use. Own-plan pairs require an exact identifier match.
Editable prompt templates
These templates are frozen into new plans or trials. Existing trials keep their original prompts. Keep the required placeholders.
Choose a plan on its own merits.
Paste each response and save a version before judging it. The latest version per planner is used in the matrix. A compromise is your manually combined plan.
Add a plan version
Spend runs on a useful question.
Use one strategy or combine them. The same exact input + builder is one trial, even when several strategies request it.
At most 64 cells per matrix and 128 saved trials. Appending never replaces results. Changed brief, criteria, plan version, builder or building template creates a distinct input. Export before starting a separate workspace.
One run at a time.
What does the evidence support?
Groups below hold the exact brief, criteria, plan version, building prompt and judging template constant. Builder configuration is the variable. Own-plan comparisons usually span groups and change both planner and builder.
One run per exact input and builder. These are descriptive human ratings, not estimates of a model’s general ability. “Preferred” is your judgement, not a win rate. Missing scores are never zero; criterion means show their own denominator.
How this workspace keeps its promises
No model calls, accounts or keys. Your work stays in this browser’s local storage; export JSON for a portable backup. Storage can fail or be cleared. The status above tells you whether the latest change was saved. Links are recorded as text and never fetched.
Plan judgement and build judgement are separate. Saved plans and trials retain exact inputs; editing the brief does not rewrite old evidence. Judge prompts omit model identifiers, but pasted content can still identify its author. Read the actual output and record evidence before a preference.
Limits: one workspace, 60 plan versions, 128 trials, 16 MB imports. No automatic execution, repetitions, statistical significance, or global model ranking.