@matthew_fitzgerald: Email: [email protected] #ai #chatgpt #engineering #Science #math

Matthew Fitzgerald
Matthew Fitzgerald
Open In TikTok:
Region: US
Saturday 24 January 2026 18:24:14 GMT
5837
247
32
10

Music

Download

Comments

precisionpulsecapital
PrecisionPulseCapital :
I used Claude and grok to double check each others work. They came to the same answer but rounded differently 5.41 and 5.40. Grok came to 5.41 and criticized Claude’s work. They didn’t know they were being evaluated against each other. I’ll send the result. Took 3 prompts each. 1 to solve the problem, 1 to check each others work. 3 to evaluate the opposite AI’s evaluation of the work and a chance to self evaluate.
2026-01-25 21:07:46
4
__hyphy__
__Hyphy__ :
Said this on the other vid too, I’m not the biggest fan of ai but…I believe if you asked the models to provide the step by step solution to solving each problem WITHOUT computing the solution - they would have performed 10-15% better. The models get tripped up when performing calculator-like tasks
2026-01-25 21:40:35
1
deadringatl
EpicShannon :
You guys are cracking me up. The fact that AI can score a 77 is unbelievable and in about a year it probably scores 95. Stay in denial lol
2026-02-24 01:48:27
0
merkin.man
Mavrick :
Will be very interested to see how this pans out. You could frame this as glazers vs haters, but it feels more people who see AI getting a C+ on a physics exam as proof of their suspicions vs people asking "okay, how can we bring it up a letter grade or two?”
2026-01-26 19:44:15
0
bra_vo12rj
RJ :
my Gemini always gets it 100% maybe ive trained it enough, I use the pro version
2026-01-25 05:51:24
2
naviwinn_
Naviwinn_ :
Any model that’s note hopelessly outdated would get a 100% on the exam easily. Current ChatGPT can do novel number theory work lmao this is nothing
2026-01-25 18:58:57
0
tzvijudah
Tzvi :
This is a very awesome way to do this. Very clever, I look forward to your results.
2026-01-25 14:02:22
1
notasaint100
NotASaint :
Could we have a student try the exam open book to simulate what a human can do with all available information, like how all these models can look up stuff?
2026-01-24 22:59:59
2
dryguy__
DryGuy | ChemE | Law :
Awesome idea I may give it a try!
2026-01-24 23:18:07
3
thesicilianisawake
Giorgio :
Just sent you the result of my tests. I used both basic and advanced prompting techniques. I'm really curious to find out what worked and what didn't.
2026-01-25 03:03:16
2
1f234
Porter Kramer :
Great idea
2026-01-24 18:49:54
1
redderif
Red🇨🇦 :
This is really interesting. I was trying to understand some PCM properties (I’m not trained for any of this) and split my question up into parts and asked multiple LLM components of the question. I ended up with wildly different numbers from at least two of them. I ended up providing the numbers I was being given to a third LLM and asked it why I was getting different values and it identified that one of the models was using a incorrect value for the density of paraffin wax. The confidence of the answer from a LLM is almost hilarious at times. I honestly find them useful to explore a concept. It’s like having a conversation with Wikipedia but you need to remember that Wikipedia can be incorrect either through a factual issue or by simply having incomplete information. I think I would approach your challenge by asking a second model to review the answer(s) and identify any issues and either suggest prompting to have the original answer be improved and/or rewrite responses with explanation of why it deviated from original answer.
2026-01-25 11:30:52
0
reubins
Rubens :
This is why fine tuning, steering, and prompt engineering is so important for those who use it in the professional world. Will be interesting to see what novice approaches have the best outcomes
2026-01-26 18:21:06
0
splatter85
Rick Kaplan :
You might also want to run the exams multiple times because AI is a probability machine. It should technically give the same answer every time based on its weights but there’s the possibility of response drift.
2026-01-24 19:31:44
3
athion
Athion :
I sent the grok answer. It thought for 6m 27s … I’ve never seen it think that long before lol :)
2026-01-24 22:38:47
5
miguerim
Michał Piotrak :
Sad the experiment ended. I was gonna use one of those specialized AIs for physics
2026-02-02 03:20:29
0
techstack101
Show your tech stack-orcal ai :
Do u use foam?
2026-01-24 23:56:26
0
britnoellej
B :
Nerd 😉🤓
2026-01-24 18:53:06
3
matthew_fitzgerald
Matthew Fitzgerald :
Just as an update. Gotten some interesting results as of Jan. 27th. Hoping to post the results this upcoming weekend.
2026-01-27 21:16:50
1
simmyball
Simmy :
What happens if you give the models notes, slides, etc? And limit the response to just using the documents.
2026-01-28 15:18:31
0
To see more videos from user @matthew_fitzgerald, please go to the Tikwm homepage.

Other Videos


About