<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[为什么模型评测总是不靠谱？说点我的看法]]></title><description><![CDATA[<p dir="auto">现在各种模型跑分满天飞，但真用起来和榜单差距挺大。我觉得几个原因：</p>
<ol>
<li>刷榜：有些厂商对着榜单优化，换个题就露馅</li>
<li>题目泄露：训练数据里可能混进了评测题</li>
<li>场景不对：榜单是通用题，你的场景可能是垂直的</li>
</ol>
<p dir="auto">所以我一直建议：别迷信榜单，拿你自己的真实任务，小样本测一下，比看一万个跑分都准。</p>
]]></description><link>https://bbs.zapai.cc/topic/85/为什么模型评测总是不靠谱-说点我的看法</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 01:17:27 GMT</lastBuildDate><atom:link href="https://bbs.zapai.cc/topic/85.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 31 Aug 2026 02:09:06 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 为什么模型评测总是不靠谱？说点我的看法 on Fri, 04 Sep 2026 14:43:37 GMT]]></title><description><![CDATA[<p dir="auto">本地跑的话显存吃多少</p>
]]></description><link>https://bbs.zapai.cc/post/301</link><guid isPermaLink="true">https://bbs.zapai.cc/post/301</guid><dc:creator><![CDATA[知识付费韭菜]]></dc:creator><pubDate>Fri, 04 Sep 2026 14:43:37 GMT</pubDate></item><item><title><![CDATA[Reply to 为什么模型评测总是不靠谱？说点我的看法 on Wed, 02 Sep 2026 02:39:42 GMT]]></title><description><![CDATA[<p dir="auto">有没有量化版，显存不太够</p>
]]></description><link>https://bbs.zapai.cc/post/253</link><guid isPermaLink="true">https://bbs.zapai.cc/post/253</guid><dc:creator><![CDATA[摸鱼被抓了]]></dc:creator><pubDate>Wed, 02 Sep 2026 02:39:42 GMT</pubDate></item></channel></rss>