@diy_smart_code: WANDR is Perplexity's new open-source AI research agent benchmark — and every LLM agent, including Perplexity's own, scores terribly on it. 500 real research tasks, 170,495 records to find and prove, and the best soft-F1 score is just 0.36. ---- 🚀 DYNAMOUS AI COMMUNITY Want to learn agentic coding with live daily events and workshops? Check out Dynamous AI: https://dynamous.ai/?code=646a60 Get 10% off here 👉 https://shorturl.smartcode.diy/dynamous_ai_10_percent_discount ⚡ HOSTINGER — RELIABLE HOSTING FOR YOUR PROJECTS (10% OFF) Whether you're shipping a portfolio, a side project, n8n flows, or AI agents — I use Hostinger for fast, affordable VPS + web hosting. Get 10% off here 👉 https://hostinger.com/DIYSMARTCODE (Affiliate link — costs you nothing, supports the channel.) ---- What you will see in this 3-minute breakdown: - WANDR: the open-source benchmark that breaks deep + wide research agents - Why 500 real tasks (due diligence, literature review, market analysis, talent sourcing) = 170,495 records to dig up and prove - The grading trick: no fixed answer key — the grader re-opens every cited page and checks each claim against the live evidence - The two walls every LLM agent hits: discovery (can't find them all) and evidence (up to 68% of cited pages don't support the claim) - The leaderboard: Perplexity 0.36, Anthropic ~0.25, everyone else under 0.13 — and $5+ per task - The hard score: even the leader earns full credit on just 1 in 7 of the records a task needs - The twist: the same grading pipeline doubles as a training-data factory WANDR benchmark (Perplexity Research): https://research.perplexity.ai/articles/wandr-benchmark-evaluating-research-agents-that-must-search-wide-and-deep Perplexity on X: https://x.com/perplexity_ai A benchmark nobody can beat yet — is that the most honest benchmark out there, or just a flex from the lab that built it? Drop your take below. #WANDR #Perplexity #AIAgents #LLM #AgenticAI #AIResearch #ResearchAgents #DeepResearch #LLMBenchmarks #AIBenchmark #MachineLearning #LLMAgents #AutonomousAgents #AIEvaluation #Anthropic #OpenSource #AISearch #InformationRetrieval #AICoding #DeepLearning #ArtificialIntelligence

DIY Smart Code
DIY Smart Code
Open In TikTok:
Region: DE
Wednesday 15 July 2026 14:35:50 GMT
405
13
0
0

Music

Download

Comments

There are no more comments for this video.
To see more videos from user @diy_smart_code, please go to the Tikwm homepage.

Other Videos


About