repo·evals
· 2026-05-26 ·master@HEAD

x-mentor-skill

alchaincyf/x-mentor-skill

🛠

71 / 100Available

📝
🗺
01Signal scanning信号发现02Content acquisition内容获取03Content understanding内容理解04Topic curation选题决策05Content production内容生产06Creative assembly创意组装07Distribution & feedback分发反馈08Learning学习
📍
📍
📍
🧬

🛑
0–29
⚠️
30–49
🛠
50–79
🏭
80–100
71
🛠· 71 / 100
  • 6 claims passed, no critical failures
  • MIT / Apache / etc., installable per deployment.install_methods
  • release_pipeline=1, recently_active=True
  • EN-only or ZH-only README
  • static-only eval; live e2e pending

#1👤
#2🎯
#3🧭
#4

User question (写推文 / 选题 / 审稿 / 增长 / 诊断)用户问题(写推文 / 选题 / 审稿 / 增长 / 诊断)SKILL.md routing table (10-row scenario map)SKILL.md 路由表(10 行场景映射)A: Write tweet — 3 hooks + 1/3/1 bodyA:写推文 — 3 hook + 1/3/1B: Topic pick — 4A matrixB:选题 — 4A 矩阵C: Review — multi-layer diagnosticC:审稿 — 多层诊断D: Growth — TweepCred + stage planD:增长 — TweepCred + 阶段计划E: Account diagnosis (needs computer-use)E:账号诊断(依赖 computer-use)Loads only relevant reference markdown按需加载对应 reference markdownActionable artifact (tweet / report / plan)可执行交付物(推文 / 报告 / 计划)

npx skills add alchaincyf/x-mentor-skillany runtime supporting Agent Skills protocol (Claude Code, Codex, Cursor, OpenClaw, Hermes, Gemini CLI, OpenCode, 50+)easy
git clone https://github.com/alchaincyf/x-mentor-skill <runtime-skills-dir>any runtime, manual installeasy
  • 📡
Host LLM (Claude / GPT-5 / Gemini / etc, depends on your runtime)
Executes the SKILL.md prompts. Skill itself has no API calls.
Inherits your agent runtime's billing — skill is pure markdown.
Runtime computer-use / browser tool (optional, scenario E only)
Scrapes x.com/@username to collect 100 tweets for the account diagnosis flow.
If not available, falls back to manual CSV upload from analytics.x.com.
· 7
7
+40
+21
+5
0
+5
0

0 / 7
pass claim-001

pass claim-002

pass claim-003

skip claim-004

pass claim-005

pass claim-006

pass claim-007

input_contract
output_contract
determinism
idempotence
no_skill_callouts
failure_mode_clarity

  • evidence_completeness='partial' (not portable) → capped at 'usable'

  • only 3/4 critical claims covered

archetype: prompt-skillcore_layer_tested? Trueevidence: partialrecommended: usablefinal: usable
ceiling 1 · evidence_completeness='partial' (not portable) → capped at 'usable'

claim-001SKILL.md 结构合法criticalskill-structure· pass
claim-002内部引用都能解析criticalskill-integrity· pass
claim-003README 承诺的能力都在 SKILL.md 有对应 workflowcriticalcoverage· pass
claim-004真实 LLM 会话产出达到 SKILL 承诺的质量criticalend-to-end-llm· skip本次评测是静态 + GitHub API 审查,未在真实 Claude Code 会话激活 skill 并跑端到端 prompt。README 自报的 A/B 评分(无 skill 7/10 → 有 skill 8/10 等)是作者内部测试,不算第三方复现证据。按 prompt-skill atom 静态评测 上限处理。
claim-005风格契约在 SKILL 里有具体规则highcontract-explicitness· pass
claim-006蒸馏来源可追溯highresearch-quality· pass
claim-007仓库活跃且作者真实hightrust-signals· pass

0%
0.00s
0

# Final Verdict

## Repo

- **Name**:
- **Version tested**:
- **Date**:
- **Archetype**:
- **Layer**: (atom | molecule | compound)
- **Score**:  /100  (from `verdict_calculator.py`, not judgement)
- **Category**:  (🏭 Production-ready / 🛠 Available / ⚠️ Risky / 🛑 Don't use)
- **Tier**: (recommend ≥90 / team ≥80 / self ≥65 / try ≥50 / risky ≥30 / broken <30)

## Plain English

Two sentences max. What does the user get if they adopt this repo today, and what would make them regret it?

- Outcome if adopted:
- Regret scenario:

## Why This Score

State the user-visible outcome first, mechanism second. Lead with what the repo *does* for the user, then the evidence.

### Top 3 score drivers

What earned or cost the most points. Reference `breakdown` from the calculator output.

- +/- :
- +/- :
- +/- :

### Core outcome
What observably works end-to-end? What observably does not?

### Scenario breadth
How many real inputs has it been tested against? Which dimensions vary (platform, data shape, scale)?

### Repeatability
Same input twice → same result? Filesystem-level or only log-level?

### Failure transparency
When it fails, do you learn something actionable, or does it swallow the error?

## What Would Move The Score Up

Concrete, testable next actions in score-impact order. Not "be better" — "add X test against Y fixture showing Z (lifts ~+N)".

1. (~+N)
2. (~+N)
3. (~+N)

## Remaining Risks

Ranked. Each risk with severity + impact + mitigation if known.

| Risk | Severity | Impact | Mitigation |
|---|---|---|---|

## Related Artifacts

- Claim map:
- Plan:
- Runs:
- Verdict calculator input:
- Rendered HTML dossier: