@nate.b.jones: Our evaluations are completely terrible right now and we need practical real world ones badly #ai #llm #chatgpt #claude #story #strategy #anthropic

Nate
Nate
Open In TikTok:
Region: US
Tuesday 25 February 2025 18:07:54 GMT
34779
1992
198
96

Music

Download

Comments

absolutekongor
🇺🇸Kongor🇮🇪 :
The real world “wild examples” are much better to show what each model can do. Listening to people’s preferences and results. The lab work’s become a stat contest
2025-02-25 18:25:23
3
user1761970942475
Jim2EZ :
Recently, I was thinking about how to understand their applicability to the real world. And thought it would be really cool and insightful to test each of them using the US military ASVAB test.
2025-02-26 02:49:33
2
tfhughes714
WTF2023 :
lol. I had to ask Claude 3.7 today to explain the difference between the ChatGPTs and the Claude’s, and why 3.7 is better.
2025-02-26 06:23:31
1
keshtfe
Kesh :
Using Upwork and real world problems is the way to go.
2025-02-25 19:29:38
2
beecee799
BeeCee :
My chatGPt has a deep research mode today
2025-02-25 22:05:46
2
prompt.jockey
Prompt Jockey :
Every time I ask for super complex prompts to test my AI against Grok, it has me do intergalactic civilization issues, or cases around military drones.
2025-02-26 18:30:21
1
chrisklaus1
Chris Klaus :
you should build your own public benchmark! 👍
2025-02-26 12:29:45
1
freudenjunge
Patrick :
I look at models too
2025-02-26 09:11:44
1
chaosz911
Chaosz :
I have tried coding with Gemini and it broke my scripts every change I made and was unable to fix what it broke. Chatgpt works a lot better. It does sometimes break my code, but can fix it too.
2025-02-26 06:37:01
2
pulau162
Pulau :
Claude 3.5 is still the coding king for me.
2025-02-26 09:56:07
2
lagudafuad
Tosin :
I can’t agree more. I jump around ChatGPT (mostly), Claude(coding) and DeepSeek. Depending on what I need to get done. The others don’t just get what or want, or just go overboard.
2025-02-26 18:20:07
1
drbolt45
drbolt :
what is best for web development for Shopify need separate HTML and CSS.. I have been using Chat GPT
2025-02-26 11:28:47
1
sarainwondertech
Sara ~ AI | Tech | Startup :
I think you nailed the problem. The pace of release also doesn’t help us users figure out “is this for me?”, so Claude taking a different route is a clear differentiation
2025-02-25 21:13:34
3
johnlesieur
John LeSieur :
Nate, the only TikTok channel where I hit the heart before watching the clip. Thumbs up!
2025-02-25 22:42:47
2
_marvharris
Marv Harris :
Experiment with what you want to be done and decide around that. I like to say does it work where and how you work.
2025-02-26 14:37:16
1
dominicwaite
Dominic Waite :
Love this.
2025-02-26 00:49:12
1
lori.in.oz
Tiggy :
I want to like grok3 like they say but no. Claude 3.7 and o1pro or deep research for pretty much everything.
2025-02-26 09:42:16
1
david_smk
David H :
Switched to Claude 3.7 for cursor. Didn’t impress me
2025-02-26 20:25:38
1
_kekius
ミ★ 𝘚𝘱𝘰𝘯𝘨𝘦 𝘉𝘰𝘣 🐸 ★彡 :
can't wait for the comet browser
2025-02-25 18:10:54
2
cmurraysr
Colleen Murray :
I understand about 15% of your posts yet I never miss one! Thank you!
2025-04-10 13:05:27
0
thewokeslowpoke
Thewokeslowpoke :
You should ask Claude to figure out the Vibe benchmark
2025-03-28 05:53:05
0
persona.non.sequitur
persona.non.sequitur :
True. For me this case is everything and differentiating between the different models is annoying.
2025-03-12 09:33:01
0
navybrattx
Navy Brat :
I found Grow to be quite amazing. I use them both and I haven’t quite figured out why I’ll use one over the other one, but Grow seems to have more complex programming abilities
2025-03-29 00:41:21
0
user67547082432398
user67547082432398 :
100% agree when all the different models came out with hype claude was always better but hype train lies. Real world high level company coding different level to the others and gap is now larger
2025-02-25 23:59:19
0
To see more videos from user @nate.b.jones, please go to the Tikwm homepage.

Other Videos


About