10xAI Wiki
Search
搜索
暗色模式
亮色模式
阅读模式
导航
此标签下有4条笔记。
2026年8月11日
基准失效的三种模式:题目缺陷、数据污染、测量工具不可靠
analysis
benchmark-failure
data-contamination
llm-as-judge
expert-audit
eval-validity
2026年8月11日
LLM-as-Judge
topic
llm-as-judge
evaluation
evals
2026年8月07日
LLM 评估
llm-eval
benchmark
llm-as-judge
agentops
evaluation-platform
rag
2026年7月25日
联网搜索评测
web-search-eval
llm-eval
ai-search
rag
llm-as-judge
citation