Prompt Compass / PROMPT ENGINEERING · EVALUATION
Keep the job intact when the prompt changes.
I built a prompt-review workflow that checks what an instruction must preserve before judging whether a rewrite sounds better.
2 min overview · optional detail below
Prompt review
Better wording. Same requirements.
A rewrite must keep the task intact. Designed test example.
THE CHALLENGE
A shorter or smoother prompt can quietly remove requirements that the answer still needs to satisfy.
MY CONTRIBUTION
I built Prompt Compass and designed explicit checks for the task contract: the instructions a revision must preserve.
01 / Workflow decision
Make the required behavior explicit.
Before rewriting, the tool makes the job explicit: the requested output, its required fields, and what to do when confidence is low. That gives the person reviewing it something concrete to compare.
Workflow, explained
- 01
Route
Identify the kind of task.
- 02
Inspect
Find requirements and missing context.
- 03
Improve
Propose a clearer instruction.
- 04
Compare
Check what the revision preserved.
Decision notes and evidence
What this instruction asks the output to preserve
Return valid JSON only with fields category, article_id, confidence, and suggested_response. If confidence is below 0.7, set suggested_response to needs_review. Success: parser accepts the response.
| Requirement | Exact instruction | What it is for |
|---|---|---|
| Output format | Return valid JSON only | A response that a JSON parser can accept. |
| Required fields | category, article_id, confidence, and suggested_response | Keep the fields required by the receiving workflow. |
| Review condition | If confidence is below 0.7, set suggested_response to needs_review. | Express a specific low-confidence fallback in the suggested_response field. |
| Success criterion | Success: parser accepts the response. | State the intended check on the eventual output. |
This is a test input. No generated response or parser result is shown here.
Alternatives and constraints
- Alternatives
- A shorter instruction can leave those details implicit. The fixture uses that alternative directly, so the comparison has a specific loss to detect instead of a vague preference for one writing style.
- Constraints and tradeoffs
- Explicit requirements take more space and still need review for consistency. They give the comparison something inspectable, but they are not a guarantee that the model will comply.
Evidence behind this account
The original developer-authored fixture preserves the complete instruction below. The local comparison implementation analyzes both versions and reports changes in contracts, missing blocks, and prompt-pattern findings.
02 / Workflow decision
Give a polished rewrite a test it can fail.
In this designed test, a shorter rewrite drops the JSON-only rule, four named fields, and a low-confidence fallback. The original remains the expected winner because its requirements survive.
Original instruction
The required structure stays
- JSON only
- Four named fields
- Review below 0.7 confidence
Shorter rewrite
“The best response for the user”
- No required format
- No named fields
- No review fallback
Designed test from the original local test data. The expected winner is the original instruction.
Decision notes and evidence
From a requirement to its check and result
Original instruction: Return valid JSON only with fields category, article_id, confidence, and suggested_response. If confidence is below 0.7, set suggested_response to needs_review. Success: parser accepts the response.
Rewrite being checked: Return a clean structured answer with the best response for the user.
May 9, 2026 report: before wins; four preservation regressions detected in this case.
preservation_regressions = [item for item in expected_preserved_terms if item.lower() not in after_prompt.lower()]| Requirement being checked | Actual local check | Expected result for this rewrite |
|---|---|---|
| category | Case-insensitive literal term in the after prompt | Missing; flagged as a preservation regression |
| article_id | Case-insensitive literal term in the after prompt | Missing; flagged as a preservation regression |
| confidence | Case-insensitive literal term in the after prompt | Missing; flagged as a preservation regression |
| suggested_response | Case-insensitive literal term in the after prompt | Missing; flagged as a preservation regression |
These checks inspect the prompt text. They do not run an LLM, parse an output, or exercise the confidence threshold. The JSON-only rule and review condition still require output-level testing.
Alternatives and constraints
- Alternatives
- The original fixture deliberately compares the detailed instruction with “Return a clean structured answer with the best response for the user.” That is the alternative under test, not a reported customer mistake.
- Constraints and tradeoffs
- The preservation check is simple and inspectable: it looks for each required term in the revised prompt. It can catch these dropped field names, but keeping the words would not prove that the resulting instruction preserves their meaning.
Evidence behind this account
The May 9 local report records the original as the winner and four preservation regressions in this designed example. Its recorded three-case pass is a developer evaluation result; the fixture’s human-review status remains unreviewed.
03 / Workflow decision
Keep the tool inside a useful conversation.
An original example conversation shows another part of the work: defining a prompt-generation feature through questions, requirements, and tradeoffs before implementation. I keep that discussion separate from proof that the feature shipped.
Original feature discussion, summarized
- 01
Clarify
Ask what the feature must do.
- 02
Specify
Record requirements and constraints.
- 03
Review
Resolve ambiguities before building.
Decision notes and evidence
A bug report should reach the bug-report conversation
Exact developer test input: Prompt Compass bug report: test Auto Flow
| Stage | Inspectable behavior |
|---|---|
| Request | The user asks to test Auto Flow and report a tool problem. |
| Implementation branch | route_status = "meta_task_bypassed" |
| Expected next tool | required_next_tool = None |
| Delivery | required_delivery_mode = "standard"; autoflow_suppressed = True |
| Dedicated unit assertion | assert result["required_next_tool"] is None |
This example documents local development after the April public beta. The supporting image shows a selected public conversation about a proposed feature.
Alternatives and constraints
- Alternatives
- The ordinary improvement route would analyze and rewrite the request. The implementation includes a separate meta-task branch that suppresses Auto Flow and lets the assistant handle the QA or configuration request directly.
- Constraints and tradeoffs
- The branch depends on recognizing the request correctly. It keeps this boundary explicit, but its examples are a narrow test of routing behavior rather than proof that every ambiguous request is classified correctly.
Evidence behind this account
The actual route sets required_next_tool to None. A dedicated unit test asserts that value, along with meta_task_bypassed, standard delivery, and autoflow_suppressed. The test source was inspected; it was not newly executed in this review.
WHERE IT LANDED
A released tool with inspectable prompt checks.
I released Prompt Compass as a public beta with installable packages, source downloads, and versioned releases. Later local evaluation work added explicit comparisons for preserved requirements, including the schema-loss fixture shown here.
What I learned and would test next
Clarity matters because it helps a task succeed. A useful prompt comparison needs to ask what changed in the required behavior as well as what changed in the language.
I would test the workflow on representative tasks with people using it for real work, compare completion and correction effort, and review whether its heuristic findings predict actual output failures.
Public beta.3 release, April 27, 2026; original implementation and May 2026 local evaluation fixtures. Later local features are distinct from the April release. The example is a designed regression test, not a customer incident or an independent effectiveness benchmark.
