github / Research & data
jev-benchmark
AI-assisted summaries and translations. Check original sources for context and performance claims.
该存储库为 TypeSafe AI 的 Jev 提供了可复现的基准测试,用于评估其在代理工具调用风险分类方面的表现。它通过清晰、模糊和对抗性任务场景,测试了模型的准确性、延迟以及置信度分数的可靠性。
译文摘要 · AI-assisted; check the original.
Source notes
This is a linked resource, not an independent verification of performance, cost or results. Check the original for current details.
- Collected
- 2026-09-21
- Discovered via
- awesomejev.com
前十名
Sponsor
- Wallpets 94 点击量$125
W
- Walltank 109 点击量$85
W
- CChowder 61 点击量$50
- Falconer 52 点击量$45
F - Lemonpod 92 点击量$40
L
- Menta 38 点击量$35
M - Sway 36 点击量$30
S
- Peon-Ping 32 点击量$25
P - Zeron 94 点击量$20
Z
- Guideless 27 点击量$15
G