#1
·
2026-07-05
·main@HEAD (tag v0.1.0)
Cheat on Content(网红作弊器)
XBuilderLAB/cheat-on-content
🛠
70 / 100Available
📝
🗺
📍
📍
📍
⚛
→
⚗
→
🧬
🛑
0–29
⚠️
30–49
🛠
50–79
🏭
80–100
▼
70
🛠· 70 / 100
- ✓7 claims passed, no critical failures
- ✓MIT / Apache / etc., installable per deployment.install_methods
- ✓release_pipeline_score=3 + pushed in 90-day window
- ✓multilingual_readme=true
- ⚪compound layer needs a logged scenario run
#2
#3
#4
git clone + bash install.sh (symlinks 15 sub-skills into ~/.claude/skills or ~/.codex/skills) | macOS / Linux (bash, jq, diff; Claude Code or Codex) | moderate |
Claude Code / Codex (LLM agent host)
Runs the sub-skills; provides the blind sub-agent (Task tool) + hooks
LLM token cost; blind scoring + cross-model audit multiply calls per prediction.
Playwright + Chromium (perf-data adapters)
抖音/小红书/领英/B站 复盘抓数(强 JS 渲染 + 反爬,requests/WebFetch 拿不到)
首次扫码登录,cookie 存内容项目 .auth/;几百 MB Chromium。属未验证 claim(需真实凭证)。
Second LLM via mcp__llm-chat__chat (cross-model audit / channel C)
rubric 升级时的跨模型独立审核,防止自我验证
仅在 cheat-bump 升级流程触发。
· 10
6 1 3
| +40 | |
| +15 | |
| +12 | |
| +6 | |
| -3 | |
| 0 |
7 / 10
passed claim-001
passed claim-002
untested claim-003
untested claim-004
passed claim-005
passed claim-006
passed claim-007
passed claim-008
passed claim-009
untested claim-010
input_contract | |
|---|---|
output_contract | |
determinism | |
idempotence | |
no_skill_callouts | |
failure_mode_clarity |
workflow_correctness | |
|---|---|
declared_call_graph | |
stop_conditions | |
handoff_points | |
atom_evidence | |
error_propagation | |
partial_failure_handling |
goal_achievement | |
|---|---|
direction_judgment | |
quality_judgment | |
meaningful_autonomy | |
handoff_timing | |
observed_call_graph | |
failure_recovery |
- core user-facing layer untested → capped at 'usable'
- hybrid-repo rule: archetype 'orchestrator' requires end-to-end evaluation of the user-facing layer
- evidence_completeness='partial' (not portable) → capped at 'usable'
- only 2/4 critical claims covered
archetype: orchestrator→core_layer_tested? False→evidence: partial→recommended: usable→final: usable
ceiling 1 · core user-facing layer untested → capped at 'usable'
ceiling 2 · hybrid-repo rule: archetype 'orchestrator' requires end-to-end evaluation of the user-facing layer
ceiling 3 · evidence_completeness='partial' (not portable) → capped at 'usable'
| claim-001 | 预测不可改由 harness hook 在代码层强制 | critical | integrity | ● passed | |
| claim-002 | 盲打分子 agent 上下文隔离 + 硬拒读实绩 | critical | integrity | ● passed | |
| claim-003 | “一个月百万粉”增长承诺 | critical | efficacy | ○ untested | Structurally unfalsifiable n=1 self-report; no dataset / attribution isolation exists to test. Deliberate skip. |
| claim-004 | 预测准确度随循环复利收敛(判断力 ×10) | critical | efficacy | ○ untested | Needs a months-long logged live run on a real channel to demonstrate convergence; out of scope for a static eval. Deliberate skip. |
| claim-005 | 15 子 skill + 路由表真实存在且成规模 | high | structure | ● passed | |
| claim-006 | rubric 升级 = 全量重打 + 排序一致 + 跨模型审核 | high | integrity | ● passed | |
| claim-007 | 1.0→1.4 schema 迁移链 + /cheat-migrate 真实存在 | high | maintainability | ● passed | |
| claim-008 | 营销话术被隔离在 README,工作文件保持克制 | high | honesty | ● passed | |
| claim-009 | install.sh 规范可逆,但会写入 live skill 目录且未实跑 | high | install | ◐ partial | |
| claim-010 | 平台复盘抓数 adapter(需登录态,未验证) | high | data-pull | ○ untested | Requires real platform login sessions (douyin/xhs/linkedin/bilibili); no-credentials rule forbids running it on live accounts. Deliberate skip. |
0%
0.00s
0
# Final Verdict
## Repo
- **Name**: XBuilderLAB/cheat-on-content(网红作弊器 / Cheat on Content)
- **Version tested**: main@HEAD (tag v0.1.0)
- **Date**: 2026-07-05
- **Archetype**: orchestrator
- **Layer**: compound
- **Score**: 70 /100 (from `verdict_calculator.py`, not judgement)
- **Category**: 🛠 Available / 可使用
- **Tier**: self (≥65) — 自用 OK
- **Confidence**: medium
## Plain English
- Outcome if adopted: You get a genuinely well-engineered discipline for turning content hunches into logged, blind, data-settled predictions — the machinery (immutable predictions, isolated blind scoring, audited rubric upgrades) is real and honest.
- Regret scenario: You believed the README ("1M followers in a month") and expected growth; the tool sharpens *judgment*, not reach, and only pays off after months of disciplined use + platform credentials for auto data-pull.
## Why This Score
The user gets a calibration loop whose *integrity* is enforced in code, not just promised. The score is held to Available (not Production) purely because it's a compound skill with no logged live run — not because anything is broken.
### Top 3 score drivers
- **+15 static_eval**: 6 mechanism claims confirmed by reading source (immutability hook, blind-scorer isolation, bump protocol, migration chain, hype quarantine, 15 sub-skills). Two critical efficacy claims untested (−4 within cap).
- **+12 maintainer_evidence**: release_pipeline (tag + CHANGELOG + migration registry) + recently_active + multilingual README.
- **−3 layer_bonus + hard ceiling**: compound layer, `core_layer_tested=false` → final bucket capped at `usable`, category capped at Available regardless of raw score.
### Core outcome
Observably real (static): the *enforcement* of the three integrity principles — hook-blocked immutable predictions, context-isolated blind scorer with hard-refusal list, 5-step audited rubric bump. Observably NOT shown here: that running the loop actually improves outcomes (growth or prediction accuracy) — needs a longitudinal live run.
### Scenario breadth
Zero live runs (static eval only). Built-in rubric calibrated on ONE Chinese opinion-video creator (25+ videos, per adapter README). Other content formats require the user to author their own rubric.
### Repeatability
Deterministic parts (hook, install.sh, migrations) are repeatable by inspection. The LLM core (blind scoring) is explicitly non-deterministic — the skill itself says to treat each blind score as a sample, not truth.
### Failure transparency
Strong. The immutability hook exits 1 with an actionable remediation message; the blind scorer emits refusal codes + contamination flags; install.sh fails loudly (set -euo pipefail, conflict prompts). Honest ✅/⬜ done-vs-roadmap markers.
## What Would Move The Score Up
1. (~+10 → lifts ceiling) One sandboxed end-to-end live run (init → predict → publish → retro) with a trigger-fire + blind-isolation log → sets `core_layer_tested=true`, breaks the compound cap.
2. (~+? ) Verify one perf-data adapter actually returns data with a real (throwaway) account → covers claim-010.
3. (~+? ) A logged multi-cycle convergence curve (predictions error ↓ over N samples) → begins to cover the efficacy claims 003/004.
## Remaining Risks
| Risk | Severity | Impact | Mitigation |
|---|---|---|---|
| Efficacy unproven ("1M followers"/"10× sharper") | High | User adopts for growth, gets a mirror | Frame as judgment-calibration; ignore README growth hype |
| install.sh symlinks into live ~/.claude/skills | Medium | Pollutes agent skill dir | Run in sandbox / use `--copy`; reversible via uninstall.sh |
| Adapters need platform login + anti-scraping | Medium | Data-pull fragile / rots | Manual data entry fallback; treat adapters as best-effort |
| Rubric only fits one opinion-video creator | Medium | Other formats start cold | Write own rubric from starter-rubrics/*-zero.md |
| Months-long commitment before payoff | Low-Med | Abandonment | Only adopt for a real long-term content line |
## Related Artifacts
- Claim map: `claims/claim-map.yaml`
- Plan: `plans/2026-07-05-eval-plan.md`
- Runs: none (static eval; no live run per no-credentials / compound rule)
- Verdict calculator input: derived from repo.yaml + claim-map by `render_verdict_html.py`
- Rendered HTML dossier: `verdicts/2026-07-05-verdict.html`
- Vault copy: `001_agent-os/skills/_evaluations/260705_XBuilderLAB_cheat-on-content/{dossier.html,summary.md}`