Read the English original ai benchmarkingmachine learninghardwaresoftware engineering AI 基准测试超越象棋,涵盖真实代码和硬件 2026年9月13日 Your browser does not support the audio element. Real-SWE发布了一个基准测试,运行大型语言模型在私有、生产级代码库上。该举动迫使供应商证明他们的模型可以处理真正服务中混乱、未经文档记录的脚本。[^1][^2][^3]