GUILLAUME.AI
GUILLAUME.AI
ABOUT NOW THINGS NOTES PROTOCOL NEXTUP

Model Trials

← All things

A small experiment, kept honest

Same brief. Different paths.

Collect plans, choose what to build, then compare the work. You run the models; this workspace keeps the experiment together.

Workspace files

What should good look like?

One task per workspace. Put tools, time budget, starting files and constraints in the brief so each run starts on the same footing.

1–5 criteria, up to 500 characters each. Rate each 1–5: 1 misses, 3 partly meets, 5 fully meets. Blank means unjudged.

Up to 10 per role, 200 characters each. Include the configuration you actually use. Own-plan pairs require an exact identifier match.

Editable prompt templates

These templates are frozen into new plans or trials. Existing trials keep their original prompts. Keep the required placeholders.

How this workspace keeps its promises

No model calls, accounts or keys. Your work stays in this browser’s local storage; export JSON for a portable backup. Storage can fail or be cleared. The status above tells you whether the latest change was saved. Links are recorded as text and never fetched.

Plan judgement and build judgement are separate. Saved plans and trials retain exact inputs; editing the brief does not rewrite old evidence. Judge prompts omit model identifiers, but pasted content can still identify its author. Read the actual output and record evidence before a preference.

Limits: one workspace, 60 plan versions, 128 trials, 16 MB imports. No automatic execution, repetitions, statistical significance, or global model ranking.

Replace this workspace?

This replaces your brief, saved plans and trial results, including unsaved pasted text. Export a backup first if you need to keep them.

Prompt ready

Copy the text below. You can edit it here before copying; changes here do not change saved trial inputs.