github / Research & data
jev-benchmark
AI-assisted summaries and translations. Check original sources for context and performance claims.
該儲存庫為 TypeSafe AI 的 Jev 提供了可復現的基準測試,用於評估其在代理工具調用風險分類方面的表現。它透過清晰、模糊和對抗性任務場景,測試了模型的準確性、延遲以及置信度分數的可靠性。
翻譯摘要 · AI-assisted; check the original.
Source notes
This is a linked resource, not an independent verification of performance, cost or results. Check the original for current details.
- Collected
- 2026-09-21
- Discovered via
- awesomejev.com
前十名
Sponsor
- Wallpets 94 點擊次數$125
W
- Walltank 109 點擊次數$85
W
- CChowder 61 點擊次數$50
- Falconer 52 點擊次數$45
F - Lemonpod 92 點擊次數$40
L
- Menta 38 點擊次數$35
M - Sway 36 點擊次數$30
S
- Peon-Ping 32 點擊次數$25
P - Zeron 94 點擊次數$20
Z
- Guideless 27 點擊次數$15
G