SATOSHI • NOSTR • AI CLAW • LINUX • ₿2B • OSINT | HODLER ∞/21M
@TutorialBTC
1.07K
subscribers
22.4K
photos
2.81K
videos
298
files
153K
links
#DTV
Don't trust. Verify.
Não Confie. Verifique.
#DIY
P&D 2022
🇺🇸
🇵🇹
🇪🇸
📚
DESMISTIFICANDO
#P2P
kycnot.me
#Hold
Poupança
#Node
Soberano
#Nostr
abre.ai/nostrminute
#IA
LLMs
#CLAW
Auto
#LINUX
OS
✅
OpenSource
⚠️
Translate
@NekoUpdates
Tutorialbtc.npub.pro
Download Telegram
Join
SATOSHI • NOSTR • AI CLAW • LINUX • ₿2B • OSINT | HODLER ∞/21M
1.07K subscribers
SATOSHI • NOSTR • AI CLAW • LINUX • ₿2B • OSINT | HODLER ∞/21M
#LLM
#Benchmarks
Are Broken—The Leaderboard Illusion
https://www.youtube.com/watch?v=FEvmk0xk84A
YouTube
How Companies Hack
Benchmarks
In this video, I dive into the controversy surrounding the Leaderboard Illusion paper and what it reveals about systematic flaws in LLM
benchmarks
—especially Chatbot Arena. As someone who’s followed the evolution of these leaderboards closely, I was shocked…
SATOSHI • NOSTR • AI CLAW • LINUX • ₿2B • OSINT | HODLER ∞/21M
#Benchmarks
#Psychology
#Animals
#Infants
#Artificial_general_intelligence
source
IEEE Spectrum
Are We Testing AI’s Intelligence the Wrong Way?
Why do AI systems ace
benchmarks
yet stumble in the real world? Melanie Mitchell says it’s time to rethink how we probe intelligence in machines.