kyle hobson
Menu
Selected work

Prompt Compass / PROMPT ENGINEERING · EVALUATION

Keep the job intact when the prompt changes.

I built a prompt-review workflow that checks what an instruction must preserve before judging whether a rewrite sounds better.

2 min overview · optional detail below

Prompt Compass

Prompt review

Better wording. Same requirements.

FormatJSON only
FieldsFour required fields
FallbackAsk for review below 0.7

A rewrite must keep the task intact. Designed test example.

My role
Independent creator · workflow and tool development
Timeframe
2026
Focus
Prompt contracts, heuristic analysis, version comparison, AI-host integration

Work shownPublic prerelease · later local evaluation examples

Explore Prompt CompassInspect the original example conversation

THE CHALLENGE

A shorter or smoother prompt can quietly remove requirements that the answer still needs to satisfy.

MY CONTRIBUTION

I built Prompt Compass and designed explicit checks for the task contract: the instructions a revision must preserve.

01 / Workflow decision

Make the required behavior explicit.

Before rewriting, the tool makes the job explicit: the requested output, its required fields, and what to do when confidence is low. That gives the person reviewing it something concrete to compare.

Workflow, explained

  1. 01

    Route

    Identify the kind of task.

  2. 02

    Inspect

    Find requirements and missing context.

  3. 03

    Improve

    Propose a clearer instruction.

  4. 04

    Compare

    Check what the revision preserved.

Decision notes and evidence

What this instruction asks the output to preserve

Return valid JSON only with fields category, article_id, confidence, and suggested_response. If confidence is below 0.7, set suggested_response to needs_review. Success: parser accepts the response.

Exact instruction from a May 2026 local regression fixture. Explanations are portfolio interpretation.
RequirementExact instructionWhat it is for
Output formatReturn valid JSON onlyA response that a JSON parser can accept.
Required fieldscategory, article_id, confidence, and suggested_responseKeep the fields required by the receiving workflow.
Review conditionIf confidence is below 0.7, set suggested_response to needs_review.Express a specific low-confidence fallback in the suggested_response field.
Success criterionSuccess: parser accepts the response.State the intended check on the eventual output.

This is a test input. No generated response or parser result is shown here.

Alternatives and constraints
Alternatives
A shorter instruction can leave those details implicit. The fixture uses that alternative directly, so the comparison has a specific loss to detect instead of a vague preference for one writing style.
Constraints and tradeoffs
Explicit requirements take more space and still need review for consistency. They give the comparison something inspectable, but they are not a guarantee that the model will comply.

02 / Workflow decision

Give a polished rewrite a test it can fail.

In this designed test, a shorter rewrite drops the JSON-only rule, four named fields, and a low-confidence fallback. The original remains the expected winner because its requirements survive.

Original instruction

The required structure stays

  • JSON only
  • Four named fields
  • Review below 0.7 confidence

Shorter rewrite

“The best response for the user”

  • No required format
  • No named fields
  • No review fallback

Designed test from the original local test data. The expected winner is the original instruction.

Decision notes and evidence

From a requirement to its check and result

Original instruction: Return valid JSON only with fields category, article_id, confidence, and suggested_response. If confidence is below 0.7, set suggested_response to needs_review. Success: parser accepts the response.

Rewrite being checked: Return a clean structured answer with the best response for the user.

May 9, 2026 report: before wins; four preservation regressions detected in this case.

preservation_regressions = [item for item in expected_preserved_terms if item.lower() not in after_prompt.lower()]
Original local fixture and actual comparison check, reformatted as readable text. The recorded report was inspected, not rerun for this portfolio expansion.
Requirement being checkedActual local checkExpected result for this rewrite
categoryCase-insensitive literal term in the after promptMissing; flagged as a preservation regression
article_idCase-insensitive literal term in the after promptMissing; flagged as a preservation regression
confidenceCase-insensitive literal term in the after promptMissing; flagged as a preservation regression
suggested_responseCase-insensitive literal term in the after promptMissing; flagged as a preservation regression

These checks inspect the prompt text. They do not run an LLM, parse an output, or exercise the confidence threshold. The JSON-only rule and review condition still require output-level testing.

Alternatives and constraints
Alternatives
The original fixture deliberately compares the detailed instruction with “Return a clean structured answer with the best response for the user.” That is the alternative under test, not a reported customer mistake.
Constraints and tradeoffs
The preservation check is simple and inspectable: it looks for each required term in the revised prompt. It can catch these dropped field names, but keeping the words would not prove that the resulting instruction preserves their meaning.

03 / Workflow decision

Keep the tool inside a useful conversation.

An original example conversation shows another part of the work: defining a prompt-generation feature through questions, requirements, and tradeoffs before implementation. I keep that discussion separate from proof that the feature shipped.

Original feature discussion, summarized

  1. 01

    Clarify

    Ask what the feature must do.

  2. 02

    Specify

    Record requirements and constraints.

  3. 03

    Review

    Resolve ambiguities before building.

Decision notes and evidence

A bug report should reach the bug-report conversation

Exact developer test input: Prompt Compass bug report: test Auto Flow

Actual local meta-task routing and its dedicated unit-test assertions. Input is a developer-authored test request.
StageInspectable behavior
RequestThe user asks to test Auto Flow and report a tool problem.
Implementation branchroute_status = "meta_task_bypassed"
Expected next toolrequired_next_tool = None
Deliveryrequired_delivery_mode = "standard"; autoflow_suppressed = True
Dedicated unit assertionassert result["required_next_tool"] is None

This example documents local development after the April public beta. The supporting image shows a selected public conversation about a proposed feature.

Alternatives and constraints
Alternatives
The ordinary improvement route would analyze and rewrite the request. The implementation includes a separate meta-task branch that suppresses Auto Flow and lets the assistant handle the QA or configuration request directly.
Constraints and tradeoffs
The branch depends on recognizing the request correctly. It keeps this boundary explicit, but its examples are a narrow test of routing behavior rather than proof that every ambiguous request is classified correctly.

WHERE IT LANDED

A released tool with inspectable prompt checks.

I released Prompt Compass as a public beta with installable packages, source downloads, and versioned releases. Later local evaluation work added explicit comparisons for preserved requirements, including the schema-loss fixture shown here.

What I learned and would test next

Clarity matters because it helps a task succeed. A useful prompt comparison needs to ask what changed in the required behavior as well as what changed in the language.

I would test the workflow on representative tasks with people using it for real work, compare completion and correction effort, and review whether its heuristic findings predict actual output failures.

Public beta.3 release, April 27, 2026; original implementation and May 2026 local evaluation fixtures. Later local features are distinct from the April release. The example is a designed regression test, not a customer incident or an independent effectiveness benchmark.

KEEP EXPLORING

CLARE & AI Glossary

From collecting AI knowledge to putting it to work.

Read next
Compasses & MCPs

Give AI work a repeatable review path.

Read next
Oppia Research Compass

Make the answer traceable to the evidence.

Read next

A GOOD PLACE TO START

Let’s find the useful part of AI for your team.

For applied AI consulting, AI analysis, prompt design, enablement, and human–AI product work.

Mesa, Arizona · Open to remote collaboration LinkedIn GitHub
Open original