#1
·
2026-05-26
·master@HEAD
x-mentor-skill
alchaincyf/x-mentor-skill
🛠
71 / 100Available
📝
🗺
📍
📍
📍
⚛
→
⚗
→
🧬
🛑
0–29
⚠️
30–49
🛠
50–79
🏭
80–100
▼
71
🛠· 71 / 100
- ✓6 claims passed, no critical failures
- ✓MIT / Apache / etc., installable per deployment.install_methods
- ◐release_pipeline=1, recently_active=True
- ⚪EN-only or ZH-only README
- ⚪static-only eval; live e2e pending
#2
#3
#4
npx skills add alchaincyf/x-mentor-skill | any runtime supporting Agent Skills protocol (Claude Code, Codex, Cursor, OpenClaw, Hermes, Gemini CLI, OpenCode, 50+) | easy |
git clone https://github.com/alchaincyf/x-mentor-skill <runtime-skills-dir> | any runtime, manual install | easy |
Host LLM (Claude / GPT-5 / Gemini / etc, depends on your runtime)
Executes the SKILL.md prompts. Skill itself has no API calls.
Inherits your agent runtime's billing — skill is pure markdown.
Runtime computer-use / browser tool (optional, scenario E only)
Scrapes x.com/@username to collect 100 tweets for the account diagnosis flow.
If not available, falls back to manual CSV upload from analytics.x.com.
· 7
7
| +40 | |
| +21 | |
| +5 | |
| 0 | |
| +5 | |
| 0 |
0 / 7
pass claim-001
pass claim-002
pass claim-003
skip claim-004
pass claim-005
pass claim-006
pass claim-007
input_contract | |
|---|---|
output_contract | |
determinism | |
idempotence | |
no_skill_callouts | |
failure_mode_clarity |
- evidence_completeness='partial' (not portable) → capped at 'usable'
- only 3/4 critical claims covered
archetype: prompt-skill→core_layer_tested? True→evidence: partial→recommended: usable→final: usable
ceiling 1 · evidence_completeness='partial' (not portable) → capped at 'usable'
| claim-001 | SKILL.md 结构合法 | critical | skill-structure | · pass | |
| claim-002 | 内部引用都能解析 | critical | skill-integrity | · pass | |
| claim-003 | README 承诺的能力都在 SKILL.md 有对应 workflow | critical | coverage | · pass | |
| claim-004 | 真实 LLM 会话产出达到 SKILL 承诺的质量 | critical | end-to-end-llm | · skip | 本次评测是静态 + GitHub API 审查,未在真实 Claude Code 会话激活 skill 并跑端到端 prompt。README 自报的 A/B 评分(无 skill 7/10 → 有 skill 8/10 等)是作者内部测试,不算第三方复现证据。按 prompt-skill atom 静态评测 上限处理。 |
| claim-005 | 风格契约在 SKILL 里有具体规则 | high | contract-explicitness | · pass | |
| claim-006 | 蒸馏来源可追溯 | high | research-quality | · pass | |
| claim-007 | 仓库活跃且作者真实 | high | trust-signals | · pass |
0%
0.00s
0
# Final Verdict ## Repo - **Name**: - **Version tested**: - **Date**: - **Archetype**: - **Layer**: (atom | molecule | compound) - **Score**: /100 (from `verdict_calculator.py`, not judgement) - **Category**: (🏭 Production-ready / 🛠 Available / ⚠️ Risky / 🛑 Don't use) - **Tier**: (recommend ≥90 / team ≥80 / self ≥65 / try ≥50 / risky ≥30 / broken <30) ## Plain English Two sentences max. What does the user get if they adopt this repo today, and what would make them regret it? - Outcome if adopted: - Regret scenario: ## Why This Score State the user-visible outcome first, mechanism second. Lead with what the repo *does* for the user, then the evidence. ### Top 3 score drivers What earned or cost the most points. Reference `breakdown` from the calculator output. - +/- : - +/- : - +/- : ### Core outcome What observably works end-to-end? What observably does not? ### Scenario breadth How many real inputs has it been tested against? Which dimensions vary (platform, data shape, scale)? ### Repeatability Same input twice → same result? Filesystem-level or only log-level? ### Failure transparency When it fails, do you learn something actionable, or does it swallow the error? ## What Would Move The Score Up Concrete, testable next actions in score-impact order. Not "be better" — "add X test against Y fixture showing Z (lifts ~+N)". 1. (~+N) 2. (~+N) 3. (~+N) ## Remaining Risks Ranked. Each risk with severity + impact + mitigation if known. | Risk | Severity | Impact | Mitigation | |---|---|---|---| ## Related Artifacts - Claim map: - Plan: - Runs: - Verdict calculator input: - Rendered HTML dossier: