@nate.b.jones: Our evaluations are completely terrible right now and we need practical real world ones badly #ai #llm #chatgpt #claude #story #strategy #anthropic
Using Upwork and real world problems is the way to go.
2025-02-25 19:29:38
2
🇺🇸Kongor🇮🇪 :
The real world “wild examples” are much better to show what each model can do. Listening to people’s preferences and results. The lab work’s become a stat contest
2025-02-25 18:25:23
3
WTF2023 :
lol. I had to ask Claude 3.7 today to explain the difference between the ChatGPTs and the Claude’s, and why 3.7 is better.
2025-02-26 06:23:31
1
Jim2EZ :
Recently, I was thinking about how to understand their applicability to the real world. And thought it would be really cool and insightful to test each of them using the US military ASVAB test.
2025-02-26 02:49:33
2
Sara ~ AI | Tech | Startup :
I think you nailed the problem. The pace of release also doesn’t help us users figure out “is this for me?”, so Claude taking a different route is a clear differentiation
2025-02-25 21:13:34
3
Prompt Jockey :
Every time I ask for super complex prompts to test my AI against Grok, it has me do intergalactic civilization issues, or cases around military drones.
2025-02-26 18:30:21
1
Chaosz :
I have tried coding with Gemini and it broke my scripts every change I made and was unable to fix what it broke. Chatgpt works a lot better. It does sometimes break my code, but can fix it too.
2025-02-26 06:37:01
2
Patrick :
I look at models too
2025-02-26 09:11:44
1
Pulau :
Claude 3.5 is still the coding king for me.
2025-02-26 09:56:07
2
Chris Klaus :
you should build your own public benchmark! 👍
2025-02-26 12:29:45
1
John LeSieur :
Nate, the only TikTok channel where I hit the heart before watching the clip. Thumbs up!
2025-02-25 22:42:47
2
Tosin :
I can’t agree more. I jump around ChatGPT (mostly), Claude(coding) and DeepSeek. Depending on what I need to get done. The others don’t just get what or want, or just go overboard.
2025-02-26 18:20:07
1
Marv Harris :
Experiment with what you want to be done and decide around that. I like to say does it work where and how you work.
2025-02-26 14:37:16
1
drbolt :
what is best for web development for Shopify need separate HTML and CSS.. I have been using Chat GPT
2025-02-26 11:28:47
1
Dominic Waite :
Love this.
2025-02-26 00:49:12
1
David H :
Switched to Claude 3.7 for cursor. Didn’t impress me
2025-02-26 20:25:38
1
Tiggy :
I want to like grok3 like they say but no. Claude 3.7 and o1pro or deep research for pretty much everything.
2025-02-26 09:42:16
1
ミ★ 𝘚𝘱𝘰𝘯𝘨𝘦 𝘉𝘰𝘣 🐸 ★彡 :
can't wait for the comet browser
2025-02-25 18:10:54
2
Colleen Murray :
I understand about 15% of your posts yet I never miss one! Thank you!
2025-04-10 13:05:27
0
Thewokeslowpoke :
You should ask Claude to figure out the Vibe benchmark
2025-03-28 05:53:05
0
persona.non.sequitur :
True. For me this case is everything and differentiating between the different models is annoying.
2025-03-12 09:33:01
0
Navy Brat :
I found Grow to be quite amazing. I use them both and I haven’t quite figured out why I’ll use one over the other one, but Grow seems to have more complex programming abilities
2025-03-29 00:41:21
0
user67547082432398 :
100% agree when all the different models came out with hype claude was always better but hype train lies. Real world high level company coding different level to the others and gap is now larger
2025-02-25 23:59:19
0
To see more videos from user @nate.b.jones, please go to the Tikwm
homepage.