10xAI Wiki
Search
搜索
暗色模式
亮色模式
阅读模式
导航
此标签下有9条笔记。
2026年8月12日
Agent 可信执行三层栈:可学习、可约束、可追责
ai-agent
security
auditability
evaluation
2026年8月12日
Agent 评测基准
topic
agent-benchmark
evaluation
2026年8月12日
AI Agent Evaluation
ai-agent
evaluation
testing
observability
benchmark
harness-engineering
2026年8月12日
AI Agent 自我改进(Self-Improvement / RSI)
ai-agent
self-improvement
rsi
evaluation
reward-hacking
alignment
2026年8月11日
LLM-as-Judge
topic
llm-as-judge
evaluation
evals
2026年8月11日
RAG 评测数据集(RAG Evaluation Datasets & Benchmarks)
dataset
rag
evaluation
benchmark
llm-judge
regression
2026年8月07日
AI 模型竞争(AI Model Competition)
ai-models
competition
evaluation
markets
2026年7月25日
视频问答数据集(Video QA Evaluation Datasets)
dataset
video-qa
benchmark
evaluation
multimodal
chinese-nlp
2026年7月20日
CursorBench
benchmark
evaluation
ai-coding
cursor
agentic