<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[为什么模型评测总是不靠谱？说点我的看法]]></title><description><![CDATA[<p dir="auto">现在各种模型跑分满天飞，但真用起来和榜单差距挺大。我觉得几个原因：</p>
<ol>
<li>刷榜：有些厂商对着榜单优化，换个题就露馅</li>
<li>题目泄露：训练数据里可能混进了评测题</li>
<li>场景不对：榜单是通用题，你的场景可能是垂直的</li>
</ol>
<p dir="auto">所以我一直建议：别迷信榜单，拿你自己的真实任务，小样本测一下，比看一万个跑分都准。</p>
]]></description><link>https://bbs.zapai.cc/topic/88/为什么模型评测总是不靠谱-说点我的看法</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 01:17:40 GMT</lastBuildDate><atom:link href="https://bbs.zapai.cc/topic/88.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 01 Sep 2026 02:22:25 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 为什么模型评测总是不靠谱？说点我的看法 on Mon, 07 Sep 2026 12:32:08 GMT]]></title><description><![CDATA[<p dir="auto">和 xxx 比哪个更强</p>
]]></description><link>https://bbs.zapai.cc/post/351</link><guid isPermaLink="true">https://bbs.zapai.cc/post/351</guid><dc:creator><![CDATA[模型评测员]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:32:08 GMT</pubDate></item></channel></rss>