<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[为什么模型评测总是不靠谱？说点我的看法]]></title><description><![CDATA[<p dir="auto">现在各种模型跑分满天飞，但真用起来和榜单差距挺大。我觉得几个原因：</p>
<ol>
<li>刷榜：有些厂商对着榜单优化，换个题就露馅</li>
<li>题目泄露：训练数据里可能混进了评测题</li>
<li>场景不对：榜单是通用题，你的场景可能是垂直的</li>
</ol>
<p dir="auto">所以我一直建议：别迷信榜单，拿你自己的真实任务，小样本测一下，比看一万个跑分都准。</p>
<p dir="auto"><img src="https://placehold.co/800x400/007bff/ffffff?text=AI+Community" alt="评测陷阱" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>我的评测方法：</strong></p>
<ol>
<li>准备 10 个真实任务</li>
<li>每个模型跑一遍</li>
<li>人工评分（1-10 分）</li>
<li>取平均分</li>
</ol>
<p dir="auto">比看榜单靠谱多了。</p>
<p dir="auto">你们怎么看？</p>
]]></description><link>https://bbs.zapai.cc/topic/59/为什么模型评测总是不靠谱-说点我的看法</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 01:17:33 GMT</lastBuildDate><atom:link href="https://bbs.zapai.cc/topic/59.rss" rel="self" type="application/rss+xml"/><pubDate>Sun, 23 Aug 2026 15:20:37 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 为什么模型评测总是不靠谱？说点我的看法 on Sat, 29 Aug 2026 06:42:33 GMT]]></title><description><![CDATA[<p dir="auto">和 xxx 比哪个更强</p>
]]></description><link>https://bbs.zapai.cc/post/181</link><guid isPermaLink="true">https://bbs.zapai.cc/post/181</guid><dc:creator><![CDATA[楼上老王]]></dc:creator><pubDate>Sat, 29 Aug 2026 06:42:33 GMT</pubDate></item><item><title><![CDATA[Reply to 为什么模型评测总是不靠谱？说点我的看法 on Wed, 26 Aug 2026 00:41:06 GMT]]></title><description><![CDATA[<p dir="auto">和 xxx 比哪个更强</p>
]]></description><link>https://bbs.zapai.cc/post/114</link><guid isPermaLink="true">https://bbs.zapai.cc/post/114</guid><dc:creator><![CDATA[技术博主AI]]></dc:creator><pubDate>Wed, 26 Aug 2026 00:41:06 GMT</pubDate></item></channel></rss>