repo·evals
· 2026-06-30 ·main@HEAD (package v3.1.0, skill-v3.8.0)

impeccable

pbakaus/impeccable

🛠

68 / 100Available

🛠
🧬

🛑
0–29
⚠️
30–49
🛠
50–79
🏭
80–100
68
🛠· 68 / 100
  • 5 claims passed, no critical failures
  • MIT / Apache / etc., installable per deployment.install_methods
  • release_pipeline_score=3 + pushed in 90-day window
  • EN-only or ZH-only README
  • compound layer needs a logged scenario run

#1👤
#2🎯
#3🧭
#4

Your UI code / URL你的 UI 代码 / URL/impeccable router (23 commands)/impeccable 路由(23 命令)Deterministic detector (44 rules, no LLM)确定性 detector(44 规则,无 LLM)Polished UI + findings更好的 UI + 问题清单

npx impeccable installNode.js (Claude Code, Cursor, Codex, Gemini CLI, +10 harnesses)easy
git submodule + linked provider buildteams wanting vendored installsmoderate
  • 📡
Your existing AI agent (Claude/GPT/Gemini)
Runs the LLM-driven commands (critique, polish, craft, etc.)
Uses your agent's own model/quota. Deterministic detector + CLI need no model at all.
· 7
5 2
+40
+7
+15
+9
-3
0

5 / 7
passed claim-001

passed claim-002

passed claim-003

untested claim-101

untested claim-102

input_contract
output_contract
determinism
idempotence
no_skill_callouts
failure_mode_clarity

workflow_correctness
declared_call_graph
stop_conditions
handoff_points
atom_evidence
error_propagation
partial_failure_handling

goal_achievement
direction_judgment
quality_judgment
meaningful_autonomy
handoff_timing
observed_call_graph
failure_recovery

  • core user-facing layer untested → capped at 'usable'
  • hybrid-repo rule: archetype 'hybrid-skill' requires end-to-end evaluation of the user-facing layer
  • evidence_completeness='partial' (not portable) → capped at 'usable'

  • only 1/2 critical claims covered

archetype: hybrid-skillcore_layer_tested? Falseevidence: partialrecommended: usablefinal: usable
ceiling 1 · core user-facing layer untested → capped at 'usable'
ceiling 2 · hybrid-repo rule: archetype 'hybrid-skill' requires end-to-end evaluation of the user-facing layer
ceiling 3 · evidence_completeness='partial' (not portable) → capped at 'usable'

claim-001确定性 detector 无 API key 运行成功criticalsupport-detector● passed
claim-002规则数量与 README 声明一致或更多highsupport-detector● passed
claim-00323 命令存在且有 reference 文档highsupport-skill-integrity● passed
claim-004SKILL.md 相对引用完整mediumsupport-skill-integrity● passed
claim-005跨 harness 安装mediumsupport-install● passed
claim-101真实 LLM 会话产出达到质量契约criticalcore-llm○ untested需要真实多 harness LLM 会话生成 UI 并逐项打分;本次为静态 + 单次确定性 detector 评估,未做 live design-generation run。compound 层天花板因此未解锁。
claim-102Anti-slop 规则在真实产出中被遵守highcore-llm○ untested同 claim-101,依赖 live run。

0%
0.00s
0

# Final Verdict

## Repo

- **Name**:
- **Version tested**:
- **Date**:
- **Archetype**:
- **Layer**: (atom | molecule | compound)
- **Score**:  /100  (from `verdict_calculator.py`, not judgement)
- **Category**:  (🏭 Production-ready / 🛠 Available / ⚠️ Risky / 🛑 Don't use)
- **Tier**: (recommend ≥90 / team ≥80 / self ≥65 / try ≥50 / risky ≥30 / broken <30)

## Plain English

Two sentences max. What does the user get if they adopt this repo today, and what would make them regret it?

- Outcome if adopted:
- Regret scenario:

## Why This Score

State the user-visible outcome first, mechanism second. Lead with what the repo *does* for the user, then the evidence.

### Top 3 score drivers

What earned or cost the most points. Reference `breakdown` from the calculator output.

- +/- :
- +/- :
- +/- :

### Core outcome
What observably works end-to-end? What observably does not?

### Scenario breadth
How many real inputs has it been tested against? Which dimensions vary (platform, data shape, scale)?

### Repeatability
Same input twice → same result? Filesystem-level or only log-level?

### Failure transparency
When it fails, do you learn something actionable, or does it swallow the error?

## What Would Move The Score Up

Concrete, testable next actions in score-impact order. Not "be better" — "add X test against Y fixture showing Z (lifts ~+N)".

1. (~+N)
2. (~+N)
3. (~+N)

## Remaining Risks

Ranked. Each risk with severity + impact + mitigation if known.

| Risk | Severity | Impact | Mitigation |
|---|---|---|---|

## Related Artifacts

- Claim map:
- Plan:
- Runs:
- Verdict calculator input:
- Rendered HTML dossier: