repo·evals
· 2026-05-18 ·main@HEAD (pushed 2026-05-15)

Claude for Financial Services

anthropics/financial-services

🛠

77 / 100Available

💹
🧬

🛑
0–29
⚠️
30–49
🛠
50–79
🏭
80–100
77
🛠· 77 / 100
  • 8 claims passed, no critical failures
  • MIT / Apache / etc., installable per deployment.install_methods
  • release_pipeline_score=2 + pushed in 90-day window
  • EN-only or ZH-only README
  • compound layer needs a logged scenario run

#1👤
#2🎯
#3🧭
#4

claude plugin marketplace add anthropics/financial-servicesClaude Code CLI (any OS)easy
Cowork Settings → Plugins → paste repo URLClaude Cowork (browser)easy
scripts/deploy-managed-agent.sh <agent-slug>Anthropic Managed Agents APImoderate
  • 🌐
Daloopa
Fundamentals data MCP
Enterprise subscription required
Morningstar
Investment research data MCP
Subscription required
S&P Global (Kensho)
Capital IQ data MCP
S&P Capital IQ subscription
FactSet
Market data MCP
FactSet subscription
Moody's
Credit & ratings data MCP
Moody's subscription
PitchBook
Private market deal data MCP
Premium PitchBook subscription
LSEG
Bond / rates / FX / vol analytics MCP
LSEG subscription
Anthropic API
Underlying model + Managed Agents API
Per-token billing + Managed Agents preview access
Microsoft 365 (optional)
Run Claude inside Excel / PowerPoint / Word / Outlook
M365 tenant + Azure admin consent
· 8
8
+40
+16
+15
+9
-3
0

8 / 8
passed claim-001

passed claim-002

passed claim-003

passed claim-004

passed claim-005

input_contract
output_contract
determinism
idempotence
no_skill_callouts
failure_mode_clarity

workflow_correctness
declared_call_graph
stop_conditions
handoff_points
atom_evidence
error_propagation
partial_failure_handling

goal_achievement
direction_judgment
quality_judgment
meaningful_autonomy
handoff_timing
observed_call_graph
failure_recovery

  • core user-facing layer untested → capped at 'usable'
  • hybrid-repo rule: archetype 'orchestrator' requires end-to-end evaluation of the user-facing layer
  • evidence_completeness='partial' (not portable) → capped at 'usable'

archetype: orchestratorcore_layer_tested? Falseevidence: partialrecommended: usablefinal: usable
ceiling 1 · core user-facing layer untested → capped at 'usable'
ceiling 2 · hybrid-repo rule: archetype 'orchestrator' requires end-to-end evaluation of the user-facing layer
ceiling 3 · evidence_completeness='partial' (not portable) → capped at 'usable'

claim-001README 承诺的 10 个 agent 目录全部存在criticalshipped-artifacts● passed
claim-0027 个垂直 plugin + 2 个合作方 plugin 全部存在criticalshipped-artifacts● passed
claim-00311 个 MCP 数据连接器在 .mcp.json 内全部配齐highshipped-artifacts● passed
claim-00410 个 agent 都有对应的 Managed Agent cookbookhighshipped-artifacts● passed
claim-005deploy-managed-agent.sh 和 orchestrate.py 在仓库内highshipped-artifacts● passed
claim-006claude-for-msft-365-install/ 目录存在mediumshipped-artifacts● passed
claim-007真实 Apache-2.0 LICENSE 文件存在mediummaintainer● passed
claim-008仓库没有 git tags / GitHub releases — 跟 main 滚动mediummaintainer● passed

0%
0.00s
0

# Final Verdict

## Repo

- **Name**:
- **Version tested**:
- **Date**:
- **Archetype**:
- **Layer**: (atom | molecule | compound)
- **Score**:  /100  (from `verdict_calculator.py`, not judgement)
- **Category**:  (🏭 Production-ready / 🛠 Available / ⚠️ Risky / 🛑 Don't use)
- **Tier**: (recommend ≥90 / team ≥80 / self ≥65 / try ≥50 / risky ≥30 / broken <30)

## Plain English

Two sentences max. What does the user get if they adopt this repo today, and what would make them regret it?

- Outcome if adopted:
- Regret scenario:

## Why This Score

State the user-visible outcome first, mechanism second. Lead with what the repo *does* for the user, then the evidence.

### Top 3 score drivers

What earned or cost the most points. Reference `breakdown` from the calculator output.

- +/- :
- +/- :
- +/- :

### Core outcome
What observably works end-to-end? What observably does not?

### Scenario breadth
How many real inputs has it been tested against? Which dimensions vary (platform, data shape, scale)?

### Repeatability
Same input twice → same result? Filesystem-level or only log-level?

### Failure transparency
When it fails, do you learn something actionable, or does it swallow the error?

## What Would Move The Score Up

Concrete, testable next actions in score-impact order. Not "be better" — "add X test against Y fixture showing Z (lifts ~+N)".

1. (~+N)
2. (~+N)
3. (~+N)

## Remaining Risks

Ranked. Each risk with severity + impact + mitigation if known.

| Risk | Severity | Impact | Mitigation |
|---|---|---|---|

## Related Artifacts

- Claim map:
- Plan:
- Runs:
- Verdict calculator input:
- Rendered HTML dossier: