Skip to main content
PodcastJuly 11, 20251:25:44

#10 Context Engineering, Grok 4, Perplexity Browser, Webflow Interaction 3

Talking Points

  • Pre-amble
  • [HEADLINE] Sakana AI introduces TreeQuest Algorithm
  • [DISCUSSION] Context Engineering
  • [HEADLINE] Webflow Launch IX3 Interactions powered by GSAP
  • [HEADLINE] Perplexity launch Comet browser
  • [HEADLINE] Grok 4 launched

Transcript

Ah, someone messaged me last night on one of my videos. I left a premium account on X is required to create live broadcast. Oh, so you need you need a live you need a premium account anyway. So, let's just remove that. Remove this. We are live. Cool. Good. Good. Good. Good. Good. Right. We can't go live with your account unfortunately. So, um you're you're you're a commoner. You are not paying for Twitter. I am. Oh, it says you need a premium account and it's I have the low tier. So do I. That's weird. Maybe May maybe may Okay, maybe it's the pro Maybe we need a different tier of Streamyard. Maybe we need, you know.

Um, yeah, we just need a higher. Maybe it was Streamyard. I didn't read properly. Right. Are we live on Twitter? Are we live on YouTube? Just double checking. Okay. Yep. Looks like we're live on YouTube. We are on a new location and I've I've tried this is going to be a probably a bit of a a bit of a rough time um because I've messaged people on my personal YouTube. You might may or may not have known it was going out on my personal YouTube to like utilize the following we already have. Um but we have moved. If you are on Twitter, please this time, I know I say it every week, but if you come over to YouTube, just subscribe.

Just subscribe and then go back to Twitter, you know. Um, just so we know we're making some headway with some subscribers, but it's yeah, it's going to be a rough old time building up the the uh what do you want to call it, but the good news is you don't get ads, so because we're not monetized yet. Yeah. Um, I don't know. Someone someone left a comment on one of the videos saying we you you should introduce yourselves next time. And I don't know whether they they were meant that we should introduce ourselves like at the beginning of every episode or like I don't I don't really understand their comment like what you because I don't think that's that's normal, is it?

To like just go, "Oh, my name's Yeah. I I think people don't care about it and even like in my videos I see where I even if I slightly talk about myself and I don't want to talk about myself just for the sake of it I want to give some context like who the hell is you know talking about this topic just to give some context that it drops like the view drops people just don't people are fickle yeah they just they just like press next or they skip like the 10 seconds I I talk about myself. So, no, we are not going to introduce ourself. I don't think that's not normal, is it? I think they might have thought it was just like a oneoff video or like that's the normal format, which is just a video, but it's not.

It's a it's a stream. But anyway, um, tell you what, we're going to also have to free, we're a little bit behind on this episode, as you can probably tell, but we're also going to have to freestyle the topics, which we've only got a few this week, so it's not too bad, but are you are you ready for this intro? Are you are you feeling like you can just wing it? No, I haven't. Do you want to do your bit and then I'll just take over for the Um, I don't know like I I can you can go about it and then I say the last bit which is you know the the last one that we have on our list that it's the only one that I know about.

No, I'm talking about you you you introducing yourself. The the the extent in which we introduce ourselves. And I'm Kabaza. Yeah. Yeah. You do that bit and then I'll jump back in and do the the lists and then you can and then you can Yeah, sure. And I have it in front of me. I just opened it. So, yeah. But you got to make eye contact with the audience, you know. Right. Okay. Let's stop waffling. Let's just let's start doing this. Let me just get into character. Yeah. Welcome to Command AI, where we press Command F on the weekly No, I it up. I'm gonna do it again. I'm gonna do it again. Yeah. Where we press Command F on the internet's AI news and some Okay, I'm going to totally this up, dude.

Uh, welcome to Command AI, where we press Command F on the AI internet news and somehow still can't find our sanity. My name is Samuel Gregory. And I'm Kverzan. We are here to decode the chaos in AI web design and development. This week we have tree quest which I think will be very important when we get to the last announcement but it's an interesting new way to basically get more out of our you know singular AI systems. The we got the topic of context engineering. So if you're into that then I'm going to riff off of that topic because I'm exploring it recently. Perplexity have decided to release a new browser called Comet, which I think is replaced Dia.

I don't know, not quite. It's quite expensive. Uh, and then finally, we have Gro 4 announced the best model. I could not see this day coming. You hear that? It's the best model. No question. Most intelligent. I yeah I didn't think that Elon Musk actually can pull it off because he's just so busy with so many things. He's doing you know decent job at everything. Well, I mean with basics they are crushing it. But is he I'm not super I don't know like the promises that he makes are big, right? He makes like crazy promises. But if you strip away, you know, what he has done from his promises and also from his political takes, he's been doing pretty good, right?

And when it comes to now AI, it felt like he started XAI just to have a beef with some altman just to stick it up to him. But now they're actually beating everybody else in the space in everything. So, he's got the money. He just bought his way into it. But we'll get into all that on this week's We'll talk about it. Yeah. Cool. So, to bookend this bad boy, we're getting quite lucky with this recently. So, I don't know that you remember uh Sakano or Saka or whatever they're called. Um but they were one of the first people or one of the first companies to introduce the idea of selfarning AIs. I think they were the first I think it was called o something or I don't know but they've introduced something called tree quest which is more important than you think and it's become very very relevant.

It basically is an open-source technique I believe um and it enables multiple artificial intelligence models to work together on complex problems. Now this is very different from agents which I think you summarized beautifully a couple of weeks ago when agents just do things right. Agents aren't necessarily working together. They're not um you know chucking ideas back and forth. they are just doing these tasks one at a time. They're just being scheduled. It's a complex way of scheduling an AI agent or an AI. Whereas this is multiple agents coming together and asking each other the questions and sort of formalizing their answer together. And it's using something called uh where is it now? It's called inference time scaling.

Um blah blah blah. Where's inference? Here we go. inference time scaling to coordinate multiple frontier models during problem solving. So, and I'm not sure how you've much you've you know whether you're joining the dots here a little bit, but you've got this B a technique called adaptive branch Monte Carlo tree search, right? The AB where is it? AB that again. the adaptive branching Monte Carlo tree search an algorithm that allows AI systems to perform trial and error reasoning while deciding which model should handle each step. So they're just like little ants working away, you know. Um and in testing they they did the ARC uh AGI2 benchmark. um a combination of open eyes 04, Gemini 2.5 Pro and DC's R1 models solved over 30% of problems compared to 23% for O4 Mini just working alone.

So it's kind of that extra juice. It gives gives these models that extra juice. And what I think is really interesting is that where did I see this now? it basically you can you can use different models together if this and because this is open source you're going to need to do that because no way in hell are open AI going to be like hey we've we've re re um got the tree search or tree quest um technique going and you can pull in Gemini and you can pull in Grock and this that and the other so no way they're going to know that but if you just so happen to get the uh open source version working you could probably play around with um multiple models working together which is really really interesting.

Um yeah so tree quest uh a flexible where is the API thing that I saw uh there's an AP basically just saying what I was just saying I just wanted to highlight it. Here we go. Features a flexible API that allows developers to implement custom scoring generation logic with checkpoint capabilities for longunning tasks. The framework supports both single model and multi-model deployments making it accessible to organizations with varying tech technical requirements which is just really cool and they claim the approach potent um the approach's potential to risk uh reduce hallucinations. So really groundbreaking stuff and I thought that the the whole you know self-learning mechanism was kind of you know the singularity but this is this is a whole new paradigm alto together that you know they're actually working together um and deciding amongst themselves which is the right approach.

Yeah. And it's proving so I don't want to get ahead of myself but are you joining the dots here? Are you How much how much research have you done with with Grock for I am I don't know honestly. We'll get to it. We'll get to it. But really really interesting. And what this means is I think it's going to be we are we have or basic I don't want to ruin the surprise but Grock showed you this is Grock heavy. This is super Grock heavy in a way. Yes. Yeah. Where they and you see it in their graphs and we'll get to them. you you they they do the long horizon thing and you see the the boost that it gets.

So Grock on its own is fine and then you've got the time context window or something they call it. Then you get that boost and this is exactly what this is. So what that tells me is that Grock are not going to be the first people to implement this and we're going to we're going to see this because it's open source. Why wouldn't other models do this by getting that extra thing and then getting that extra buck because we know the price of Grock. uh we're going to start to see a new tier of our uh AI chat bots that offer this agent, you know, working together tree quest functionality. I love that tree quest because they're all branching off and doing their thing.

It's cool. Yeah. Um and something that uh let me just actually maybe just add this here. I wasn't sure. Okay, let me just do it. So, does that make sense? By the way, I want, you know, I'm thinking about the edit by the way. I'm thinking about that. Does that make sense? Do you follow that just a lot of information? Yeah, that that so what I'm going to show is something that is relate kind kind of it's not related actually. It's a different way. Um, it's a another tool called round table. Should I just just share it and explain? go for it in the edit. If if it's not related, then we take it out.

I want to do it that we don't have to stop my because it's going to do my screen share now, isn't it? That's the problem. Um, yeah. So, you continue. If I remove then that's just going to mess with the edit a little bit. So, let's just kick this into gear. Cool. What's going on there? Do you want to zoom in a bit? Yeah. Uh, on the wrong window. Let me zoom in here. Oh, I love this. I love dark green like this. I love the It's like perplex perplexities Brandon. All right. So, we have this new tool in the town in town. In the town? No. Yeah. Whatever. So, we have this new tool in the in town.

Uh it's called roundtable.now. And it's essentially just a bunch of different AIs around the table discussing something together. So I just signed up literally right now and I asked is there such a thing as objective morality? So the idea is that each uh GPT each um AI tunes in and they talk about you know what you ask and the idea is that they discuss this um like between each other and then they should come to a conclusion. So here I'm saying like conclude and then we see each essentially chatbot uh trying to come to a conclusion and they see each other's responses. That's also kind of cool to you know see to ask the same question actually but in one place and then you get your answers from all these these different chats and the point is that they also see the the answers from each other.

So they can actually continue talking about it until they reach I don't know conclusion or you could tell them disagree with each other. So you I don't know. You'd probably see them fight uh over what's you know true burn a bunch of trees just to see some AI chatbots fight with each other. That's the goal. Bloody Europeans and your burning of trees. Yeah. And uh yeah this this is uh funny because this is also going to be pretty expensive I guess. And and what I'm seeing there you've got GP GPT4 search. If you scroll down, you've got Deep Seek. Yeah, that's it's up. But yeah, again, Deep Seek and then Gemini 2.5. So, are these are these models are there different models working together one by one?

Passing the buck? Yeah. Yeah. Um, very cool. Very cool. And I don't want to get too sidetracked, but how much was this? You just bought this or was it free? It's um on the free tier and I On the free tier, I believe. Okay. They actually have 60 bucks. 60 buckaroo is quite expensive. Um, so I have five models right now. Lovely. But like if you paid unlimited, you get 20 models at once and then it's a,000 a month or a thousand a year? A month. I mean, you know, this is all starting to come into play. So here's here's here's my take on it. Right. First of all, this is pro is this open source or is this closed source?

No, this is like a very new tool uh that I found yesterday. Okay, fine. So, yeah. Well, it'd be interesting to know if it's open source. The other thing I think it's going to achieve the same thing, but it's just done in a different way. So the differentiator comes down to it whether it's open source or not because they're going to keep their own you know round table are going to hold their own um try and keep their would you call it the subscribers they're going to want to try and earn from subscribers whereas Treequest is sort of open source and that means Gemini and open and AI and all the rest of it are going to adopt that and make it their own sort of thing.

So there's an incentive there just to keep it internal. Whether it's using tree tree search or tree quest, I'm I don't know. It'd be interesting to know, but it's uh probably probably the same result just maybe achieved in a slight different way. So but it's um but yeah, and you can see the the impact on price as well. What this does, you know, really I I believe this does the entire thing much different from what you showed. Uh but the results might be very similar but the way they are doing it is very different. Are there any benchmarks on this? No, I don't think so. This is literally just taking the APIs from different uh chat bots and just putting them together.

Yeah. Yeah. So I could I could add this to Jupyter chat sort of thing. It's just uh very expensive way to do it because you're just adding, you know, more context more context exponentially, right? exponentially because for each one if you want them to see each other you'd you'd have a a tons of like input and output just uh running in the background. Yeah. Yeah. No, it's really cool. It's really cool that we're seeing still seeing a few little advancements here or there. Um, and I would love to see some sort of benchmark or some sort of like what the, you know, what the result is of what you've just shown as well, but probably, yeah, probably the same sort of result.

I'm just really keen to see what the opensource nature of Treequest brings to the other models because again to ruin the surprise, Grock seem to have already implemented it because this was only released this was released probably Monday, you know. Yeah, they announced it on Monday. So, uh, yeah. Um, cool. Really exciting. So, this is the part where I was going to riff off of context engineering. Have you heard about context engineering at all? No. And in fact, I I commented on someone on LinkedIn just now because they were they they put a post out on LinkedIn saying that it's proven that AI assisted coding is actually slowing developers down. Now, uh, this is from research, uh, research paper.

Maybe I can get it up while we chat. Some some sort of research paper. Um, and I responded to that because of course I would and said, you know, it'll be interesting to see what methodology those developers employed to say they're doing AI assisted development because, you know, like like a lot of people, myself included, when I approached AI assisted development, I kind of went in thinking, right, how can it help me write code? and was not was getting some help. Co-pilot is great. Um actually what it's kind of difficult now because co-pilot means so many different things, but like autocomp completion was great. It's a great little start because it was is helping me write code.

But as we start to get into the more agentic stuff, we needed a new approach to write code using AI. And um I'm just looking now traditional coding practices don't cut it. Cut it. What way? Speed. Uh, basically, yeah, I'm just reading the the message there. Maybe I'll share that later. But like, um, we need we needed new methodologies in order to utilize AI to write code. And then we kind of got into what I've been talking about and trying to understand, which is like the task master stuff where you create your product requirements document, you develop a list of tasks based on that, which is still very, very relevant. But um it l what it lacked was a lot of what came about later on which is this idea of memories and and you've got the Gemini files and you've got the Clawude files and now the replet files and these files that you pepper throughout your codebase to sort of give it context as to what it's working on and what it should do and how it should format things.

So you're constructing a network of rules that um can help the AI write code. Yeah. And I've been experimenting with it and then this idea of context engineering is coming into play which is it is a a library uh that you can download on GitHub and just start but all it basically is is a bunch of files that tell the AI how to format stuff but as long as you reference that file you don't need a library or anything like that. You just need a bunch of files that you can reference to say implement this task, implement this feature using these rules and this rule set will say how to format things and blah blah blah.

And this is how I understand context engineering. You're also giving it documentation that it should should search as well. Especially because LLMs are so behind on documentation. I still have a hard time getting LLMs to implement Tailwind 4 libraries and I and it's because they're so behind. You know, Tailwind 4 is a relatively new um approach to Tailwind. It's still trying to spin up things in Tailwind 3. So, I need to tell it to use Tailwind 4. I need to tell it these are the libraries I'm using. I need to tell it where the documentation you get the documentation from. And there's there's um an MCP called context 7 which you can get the most up-to-date documentation for these libraries on there.

So you can sort of bypass it using its internal knowledge. And then all this to say this idea of context engineering is becoming the way to to to AI code. Not vibe it's it's a nice blend between vibe coding and AIC coding. But you're you're writing less code and you again you explained things beautifully last week where you said we're not we're we're becoming the co-pilots and we're guiding we're almost guiding the pilot, you know, and our our agents are are the the pilot in this instance and we're just guiding them and we're giving them all the information they need to be able to do what they need to do. And um I actually played around with this idea of I I want to share it but not just yet because it's not completely um complete.

But I completely AI using AI completely built my personal website using AI. That's a terrible way to phrase it, but ultimately I did the research. I I I used Manis and maybe I should make a video on this, but yeah. Um I used Manis to do research on me. So it went off to my various different um YouTube, Twitter, LinkedIn, did all this sort of stuff on my company Jupiter and the draft did all this research and pulled together a wireframe and a plan for this website and then I took that plan. I then um built with again with um context engineering built like a CMS powered website with it. I I think I've got a design that's looking pretty good because then I use Stitch.

Do you remember Stitch? Google's to design it as well and it's been brought together completely with AI assisted development and the to bring it back to context engineering the actual coding part of it was so much quicker I spent hours on stream trying to learn taskmaster and understand taskmaster and ultimately I was starting to feel like you know it's better if I just have files here or there I was doing context engineering without it being defi a lot of people were doing this I'm not saying I'm trailblazing here, but like a lot of people were doing context engineering without actually doing context engineering. So, um the fact that I did all of this in probably less than a day, probably less than a day.

Um counting up all the hours, it was a bit bit bumbling around and stuff like this, but the actual development was actually really straightforward. And I'm really really enjoying context engineering. to bring back to the comment that I I said at the beginning. It'll be interesting to see what my friend says, but I think a lot of a lot of development the way we develop has complete is completely shifting and we are we should be writing less and less code really. We need to be just letting the LM worrying less about the code. sorry. Worrying less about the actual nitty-gritty of the code and just being more high level thinking, being more of a manager, being more of a shepherd, being more of more of a co-pilot to the AI.

And I've seen success and I I do I want to make a video on it on kind of context engineering in general and how I put together my website, but I've just got to launch the website first. content as well. The actual content on the the only thing that wasn't vibed. It was almost it was almost AI done was the actual my actual profile picture. I generated it, but it looked fine. I think the profile picture looked fine, but the quality of the image wasn't there. So, I just snapped a photo on this camera and then did all that. That's the only thing. When you say you use the AI to do the research, uh did you like tell it where to go or it just decided?

I I told it where to go on a few things like I just said go out and but I don't think I had to. I think I needed to give it more of a clue as to who I am. Like I couldn't say go out and research Sam Gregory because there's a bunch of Samuel Gregory's out there, right? So I probably need to give it I knew I needed to give it a few little URLs here or there to kind of contextualize it. But um I'll probably on the video go through the actual sources it looked at. But it'd be interesting to see if it did find a few little bits and bobs here or there.

Actually, it did, you know, because it found um the full stack agency, which is like an experiment I ran a few years ago, and it was like that's crazy that's brought in full sec agency stuff, but like Yeah, that's um I I needed to poke it a little bit to where to go. Yeah. Yeah. But it'll be interesting to see what the result is because again, it was just all all done through AI designs. It's I mean like you know me I'm not like we we've we've talked time and time again how we've said design is just not there because it's not it's just it's derivative right so it's not going to be groundbreaking design but as someone who like sort of has a little bit of taste you can you know decide when when he does or doesn't like a design I think it's fine you know it's getting decent it's not exciting design it's not creative I don't I don't see many creative thinking from AI in any form or shape.

It's just it will this will be the next kind of frontier that is like relevant to us when we talk about the tree quest kind of thing. It's about how they can work together to create something new. But you'll never, you know, I I'll I'll rant for you give me 15 seconds maximum to rant. But the thing is about designers is that they they want to be creative. they want to do this, but then they end up just copying whatever Apple's latest trend is anyway. Like they want this flexibility, but barely anything I see is original. Apple uh we see that what they do is not always, but very often is very creative, right? That we we should have put this in today's episode by the way, the latest update for liquid ass.

Yeah, maybe we riff on it at the end. Well, they are um whether I I saw like a few tweets that Apple is dialing down their they basically backtracked. It basically looks like the previous version because the liquid glass thing doesn't doesn't work from an accessibility standpoint. Yeah. Let's see. Um anyway, yeah. So, you guys copy liquid glass because of why? We didn't copy it because you always copy Apple. We were No, we we did we did we did it for just because of the the views. Yeah. The views. Yeah. Yeah. That that that's how we decide quality. Will it get views or not? Yeah. But the creativity of an LLM is is yet to be seen and and you won't see anything crazy new.

And maybe you don't want that though. Do you really really want that? Like is that saving you a lot of time? You know, surely that's your outlet to be creative. You don't want it to be. Well, if it's not creative, if it's just, you know, outputs average, then it's meh. Yeah. But you but that's your that's your opportunity to shine then that it puts in the the boring stuff. It puts in the fundamentals of like layout and hierarchy and and maybe some contrast. And then there's your point to be like, "Right now, I'm going to be the superhero. I'm going to waltz in fix this puppy. You know the issue is we don't have a lot of editability with LLMs when it's not, you know, code or text.

Like with images, we do have some sort of like editability, but we don't have that much of creative control. Not yet. And when it comes to design, I would love, you know, to draw something in Figma and then maybe have a style guide and Figma AI could pick that up and create like a very clean. Yeah, like a perfect design that is boring. Okay, I'm fine with that. And then me as a human with taste, I come in and, you know, add a little bit of like interest to it. So I can make it, you know, just a little bit of my own magic. Yeah. And that's fair enough. They are not there yet. And you don't want to be writing prompts to say make this border radius 12 pixels.

No, that's what I've been saying for a long time. I think this chat UI is it's something that we have for now, but it won't be forever. Like I I I don't think it's a good UI. I think we will we will create more software with like more control with more UI where you we can click things um in the directions that we want. We will have a lot more UI. I think you and I need to come together because I think what you're saying like as a technical person what you're saying doesn't make sense from an LLM standpoint but I think I'm too technically minded to to to whatever. So I think you we we need someone who's perfectly in the middle to bridge those two worlds because as a technical person I just don't see I I I don't know what you're asking.

It's like what are you asking from this LLM? What are you asking for from a lot of times when we design talking about design? Um, you know, e even if somebody is painting, you see a lot of times you correct yourself like you you try something, you step back, you look at it, it doesn't look quite right, you paint over it. In Figma, you know, you you change things, you have like sliders. Like in Figma, you literally have invisible sliders, you know, when when you go to an input field, you drag it and then you see the corner radius, you feel it, right? You see it in the design and if it feels right, it feels right.

Same with colors. Like you put things together and you just vibe with it. You know, you listen to music and you design. Um AI is just not there to help with like, you know, like with chat. AI is not helpful because I don't I don't want to say h well that corner radius does not feel right. Can you make it feel right? No. What do you want to do though? What do you want to do? I I want to have UI. I want to have literal UI controls on the fly. So if if I'm asking um an LLM to create a component for me, it's a web component. Imagine whatever that is. a shader or if it's like a model, right?

I want to have some UI control to change some properties. And there are like tons of uh libraries for that like tweak pane. Uh I don't know if you know the library or like the the one for 3JS. It's called theater.js JS where even like when you are creating a 3D scene with code it's difficult to go in the code and change you know every time the the the camera position. So what you want to do is to bind that camera position to that UI. So you you don't have to change it from the code. You just eyeball it by literally drag and dropping uh dragging um the sliders around. Sames for animation like we we can uh mention the web flow animation being out since yesterday.

That's where like you have a timeline and you can see how things how how long things are. And with just a chat interface I get I get that with a chat interface I get to a um component really really really fast. But where it sucks is when I want to make small changes because with small changes then a chat interface is much slower. It's much slower for me as a human to interact with it. If we get neural link where I can like literally think of you know talking to the LLM then that won't be the case. Maybe the chat is you know just the first step and then the next step is having some sort of UI and the next ultimate step is just to integrate all that in our neural link.

Let let's move on but get one of your fancy animators to literally animate and previs this whole thing because what you're saying doesn't make sense. Anyway, I look forward to sharing this whole, you know, this this website with you. But like um yeah, this idea of context engineering is a real the the missing piece I think personally to this whole AI assisted coding thing that um people need to scrub up on. Now let's go to an ad break um and then we'll talk briefly about web flow because it's relevant to the ad, I guess, right? Yeah. Yeah. Support for this episode comes from Flowbase. If you are building in Web Flow, Framer, or Figma, Flowbase can help you build much faster.

With over 4,000 components to choose from, you have a huge variety of wireframes and super clean, nicely designed sections that you can put together by just simply copy pasting them to your project. Over 15,000 icons as well, over a,000 illustrations. They also have a super helpful Web Flow app called Boosters. With Boosters, you can add things like sliders and mares and countups to your project with a few clicks and without any coding. So yeah, big thanks to Flowbase for supporting the show. If you want to try them out, go head to flowbase.co. And just for you guys, I want to show discount command flowbase all capitals for 20% off any plan. This is a limited offer, so get in there early.

All right, that's it. Let's continue with the show. Cool. Thanks, Flowbase. 20% off. Awesome. Um, so Web Flow Animations, they find they're rolling it out, right? Over the next month, they're going to be rolling it out. Have you got access? Um, I just checked. I don't have access. Let me see if I can share something on the screen because people ask, do you have access? No. No. I I might do, but um let me let me have a look. But it's the thing we've been after for a long long time. Um which is horizontal um yeah horizontal animation timeline thing, but it's very reminiscent of our old boy uh Pinerow, which were a kind of OG GASAP um sorry, I'm doing too many things at once right now.

OG GASAP timeline editor, but I'm really I'm really excited. Yeah, it's super cool. Like GSAP is the the standard for animations right now and you can very soon like as as you said they are rolling out create like much like easier anim like you can create more complex animations much easier indeed. Right. I'm going to especially when it comes to animating text like previously in web flow it was really difficult to animate text I mean it was easy but it was time consuming because you had to animate each letter you know letter by letter and you had to each in a span if you didn't want to code you had to do a bunch of like manual work but right now you can just you know use the GAP uh text split which all is included for free and I think it's going to help us create like much more like also better performing uh websites with a lot more animations very customized and also accessible because web flow is um actually going to tackle accessibility as well with gap there are uh they I think they talked about uh the reduce motion Oh yeah, I know.

Yeah. Yeah. Because I think that's the thing built into GAP. Have you I mean, have you seen the UI? I'm just loading it up now, but just keeps crashing. Literally, I'm on a crash loop. Yeah, the uh yeah, uh I can share my screen because I have it on the screen. So, yeah. No, I don't have uh I don't have interactions just yet, unfortunately. Okay, I'll we'll just check it from the YouTube. All right, so this is how the panel is going to look like. You might want to reload because you might have it. So, what we have here is this interaction icon. Then we have our canvas and then on the right side we have the the controls.

And in the bottom we have the timeline where we see different elements being animated like different animations. So we see here a text animation with stagger as you see we have a preloader we have logo move up and if we play this this is a YouTube video you see how this works. So you can literally drag things and play the animations. And you see how nice it is. You see how the text the text here you see how it is animated like word by word. Yeah. You saw this. So previously you had you couldn't do this uh without code and just in one go. Now you can do it in just literally one go. Here on the right, if we look at this uh closely, we have um the target.

We have the trigger, but also the target. The target is like are we targeting a class or an attribute or an ID, you know, a bunch of like different targets that we previously didn't have access to. And then here you have the properties that you can animate. So you have a bunch of like from height to size and you know to transforms but also class. So you can actually toggle add or remove a class. Very cool. Yeah, I I'm looking for something that allows you to change HTML values or attributes because this is the thing that I my video that got big a while ago was talking about how the interactions panel is not accessible.

you're if you're going to animate something in like a popup or something like that, you need to change values to to let a accessibility user know that they've just opened a modal and that the menu is open and that there's this there's a drop down and things like that. So although we can't see it from the video or whatever, um that's what I'm that's what I'm keen to see. That's what I'm I'm wanting to see. So yeah, but overall, you know, the some of these examples look pretty cool. They get they released three examples, didn't they? Um from from some people who've had early access and they've done some a few nice things, but it's uh it's you know, you expect nothing less really, do you?

Yeah, this is this is really cool. Um yeah, I think that Yeah, I think it's going to change the game again. uh especially comparing Web Flow to Framer and a bunch of other tools. Now you can not only create like richer animations but also like easier like I really want to emphasize on this point. There are a b there are a bunch of pre-made animations and you have an animation kind of like library system where you can just make an animation let's say for text and then you can apply this to other things and that is really really cool and it just works. You don't have to worry about uh you know like previously you had to match the class like a bunch of things.

Now you can just have like one class and then say whatever children that exist within that class stagger them and it just works. Yeah. Because I' I'd hate to think how long the JavaScript library that Web Flow used beforehand how long they've been how old that is. basically how whether it's still performant, whether it follows modern web standards and things like that. Whereas Gab have always kept up to date with modern web standards and they know how it's done. Whereas yeah, I'd hate to think what the the amount of junk that gets put that got um has been left in web flow with uh what do they call it? Interactions 2, whatever it is, you know.

So yeah, I ex2. So this what you see here, we did this manually. Like I went in and took each letter and literally like turned it into a span and then animated that in individually. All of these are animated individually. Painful. Right now um I could probably make this with just two like like Yeah, two different animations. Yeah. Just build the animation with a class, attach it to a class, and just apply it to those pieces of text, right? Yeah. Yeah. That's it. That's it. Um, excellent. Do do one thing for me on Twitter. See if you can join whatever it is that I'm broadcasting. Whether you can actually like join that. Yeah. Uh, I'm I'm watching the our Twitter show live.

Can you can you join that as it under your account? Yeah, I'm there with my account. I just said hi in the chat. As in like does it show do you show up as like presenting or whatever it is? You know when you can huddle and then it shows different people presenting and whatever. You can't do anything like that. No, I don't see any option for let me check notifications. No, I mean I haven't ignited you or anything. I was just wondering if you can click a button and then that happens. But um let us uh before pressing remove that I know what you mean but it's just just um don't worry about it. It's all good you and it says we have 20 plus viewers but I Hello viewers come over to us on Tw on YouTube and give this new channel a subscribe because uh we need all the help we can get to build this new channel up.

It's not on my channel anymore. Oh, you I see a fancy website. Yeah, this is a lovely little website. This is a lovely website. So, this is so perplex you might be saying bye-bye to for to for to to Diaz pretty soon if they can get their act together. So, this is all scroll based animations. I do like Perplexity's brand to be honest. It's I mean it's space and I love space and blah blah blah. So, Replexi have released a browser and the interesting thing about this is that to be honest, it doesn't offer much more than say DIA does. You've got the context. You can at uh let's see if I've got a window over here.

Perplexi's new AI browser. Um you can at mention uh we won't get into it there, but you can at mention kind of other windows. uh and get it to, you know, as you've been doing with your with your little EU thing. But the big feature here is it's got the auto agent mode, which this see this here. And if we scroll over, I think that might be his mouse actually, but he does get to work. Here we go. This this green glow that happens around the window. Now, this is the agent actually navigating and clicking around and doing things on the website and actioning whatever it is that you need to action, which is which is what DIA promised, but they haven't released it yet.

So, this is really really quite cool and I'm quite excited about this. There are still some bugs as you know if you go to to Matthew Burman's video here he does it does kind of get a little bit confused about like what it's meant to do. Sometimes it actions things sometimes it doesn't. So it's very much I don't know whether it's officially in beta but you know he he struggles to kind of get this actual blue box kind of appearing around. Um, so yeah, it's it but it's it's what we what we hope and dream and it's called Comet and it's pretty d gosh darn cool. It can do a bunch of stuff. Prepare and locate the tweet.

I want to show you something that's a bit of an annoying thing that it does. Here we go. So my comet agent posted this tweet and then it added on the end this create created with comet assistant which is really freaking annoying. Like why have they done that? Especially Especially when you come to learn who has access to this thing. Um, where is it? Perplexi comment. Where are we? Uh, I want to I want to show the pricing. Have you seen the pricing? No, this is like my first time hearing about it and I'm Okay. So, yeah. Um, I want Basically, it's 200 bucks a month to have access to this. Wait, what? Not I mean you get Perplexity Pro, but you you know it's like a it's not just the browser, but you get access to this.

I want to try and find the the uh what do you call it? The actual pricing page, but it's 200 bucks a month for this and it's still kind of buggy. Yeah, we've been lied to. We'll get into this with Gro for but we've been lied to. We were told that AI will get cheaper and it's just getting more and more expensive and every freaking company is still like losing money on it. Yeah. While it's getting more expensive. Yeah. Yeah. Well, this is why I said low key, I think we'll be looking more and more to running LLMs on our local machine. And cheers for the love. Whoever gave us some love on YouTube, can say say hello on the chat.

Uh I want to find here we go. Here's comet free. I can't believe uh perplexity max. So unfortunately they really they really um what do you call it? They abstract that away from you, don't they? Yeah. But but this is this is interesting and I'm I'm very interested in this browser the this new like era of brow browser um where Proplexity is known for being they're not the most minimalist with design. They are I would say even like maximalist. Yeah. Yeah, not in a bad way. Not in a bad way, but they tend to maximalize on the design and uh like anything that they can do they do with design, but DIA is very like it's much more minimal and I actually really like it.

But we'll see which one wins. And this agentic browsing I think is it's really cool and it might give it it will be very interesting because there was a moment where we thought okay what is going to happen to websites because people are not going to browse because chat bots will just browse in the background but now we have AI browsers and now chat bots are going to actually browse so that would interesting to see and how they how they are going to use a website because these will you know just control the cursor and click around that would be interesting. Yeah, it will be can do stuff like on your can you make these chat bots, you know, do things like you add some type of meta information to your links and make them click those links and I don't know buy stuff that probably not that be next level.

What do they call it? Not SEO be like a io aentic. Yeah. No, no. A aentic. Yeah. AI oh I don't know whatever optimization I don't know AI optimization AI search optimization cheers for subscribing Vera Patel uh DS doing a lot of improvements it really I think as as a base browser it's um I'm just showing you the price here with access to comet but um you know it's a very basic browser I'm not that demanding a oe yes um as a basic browser it's fine but like Yeah, they they they're they're building in public I think and um unfortunately I think a key feature that I'm personally after comet is answering the question. Interestingly though, this is this is hot topic, right?

Um I will find something and maybe I can show it to you. Um, I can't sources videos. No, trying to think. Okay, I'll I'll ruin the surprise. The I did I've got Perplexity Pro. I pay for it and it's got this thing called uh Perplexity Labs and I sent it off on a a bit of a I sent it off on a it's an agentic thing kind of similar to Manis and things like that. And I was watching it do its thing. And I was watching it navigate around websites, clicking into websites. Now, this to me was it's work um Comet auto browse working under the hood of I'm just trying to see where it was.

Um under the hood of Perplexity Labs, which is a pro feature. So I think they've they've baked this in. They've just kind of like abstracted it abstracted it from, you know, do you know what I mean? Like it it's they've just taken the technology and exposed it in a more, you know, way. I really would love to show whether there's some sort of recording. I've even just forgotten which task I asked it to do. You want to share the screen because like you're not I was just afraid that I might be showing something personal or whatever, but like I just can't find it. I mean, I've asked basically I I I transcribed one of my own videos and I was like, um, I thought it might be under videos here, but it's just, you know, it's just done that.

But I was just trying to see, but I was I was watching it. I was watching it. Uh, maybe we can just do it again. Let's let's edit this. See if it does its own. See if it redo it and and we can see it working. But I guess the point is is that you you can kind of get sneaky access to Perplexity Comet uh just by using labs. It's just not as clear that it's actually doing it. And to be honest, I couldn't tell it to do something like, you know, well, I actually haven't tried. I could I might try and tell it to do something specific, but yeah. So, Plexi Labs just takes too long just because it's an agent rather than a just a AI.

But yeah, I'll leave that running and we'll pop it up if if I do notice that it's doing something um you know of note. But yeah, that was just something I noticed. Yeah, and it's going to be really interesting like uh you know, not just to do browsing and finding information, but do things like most of the tools that we use nowadays are in a browser. Mhm. Like when will we get to a point that an AI and a browser DIA or comet can do stuff? I I don't know when but that would be cool like you know to ask them to do something in web flow you know just goes in and like clicks it around done I don't know like an idea or whatever else like it's it could be a Google sheet it could be using just notion you know you could you could theoretically you could run your entire notion just with voice commands right you go into dia and you say oh go manage these projects and add these I don't know like these tasks add a QA and then one when will it be when it can run a QA you know the website I'm like ah these are issues with it so you need to solve them okay go back and solve them I mean this I mean yeah these are these are great but like I'm I suppose I haven't really paid attention how much I am in the browser I am a in the browser a lot but I can't help but feel like as I said last time, I've got OBS, I've got Final Cut, I've got all of these applications that are not in the browser and they would be kind of they just this is where Warm Wind, which you spoke about in the past, uh comes to play.

Oh, we might be getting something on Perplexity here. You might be getting something I've just reran one of my things. Yeah, you see this? This this is a browser. I don't know whether you can see it. It's very f Here we go. Yeah. So, it's gone onto our thing. And let me click on the more button. So it looks look see this is exactly what I was talking about. So this is comet under the hood. Yeah. And it's it's kind of scary, man. Let me click. Let me click. So and you saw something come up here. Maybe we can open that. Let's see what it's doing. Then maybe we can watch it do its thing.

How crazy would that be? This is our YouTube channel, by the way, guys. Go subscribe to it. Um yeah. So I can't it's not we're not seeing some live stuff going on here but yeah but yeah this is it. It's opening stuff. It's it's opened this page. It's not just visiting. Look let me click on the more button. That's what I'm talking about task. Are you using um DIA or No, this is Zen. So this is nothing to do with the browser. This is Yeah, I know. But like are are you using DIA in and are you hoping to like get I'm using DIA as my dayto-day browser. Yes. For our streams like this, I use Zen just because uh it has the full screen thing so it looks a little bit nicer.

But for my day-to-day, yeah, everything is open in in Dia. Yeah, Dia. I hope that they will add this as well because they have the vertical tab now, but they don't let you close it completely and, you know, have this fullon experience, which I think they should and they probably will add it. One thing that I uh that I started doing uh opening my WhatsApp in Dia. I'm not using WhatsApp in as an app anymore. I'm opening it in dia because I'm uh like I have two two clients that communicate with me in German and sometimes I just don't feel like texting in in German. So, I just, yeah, write it in English and Dia does it in German and I just click insert.

Yeah. Yeah. I need to I I You know, you should make a video like top five ways I'm using Dia. Yeah. I'm um I'm going to make one. I'm still like transitioning. I'm still transitioning. But yeah, Dia is becoming my daily and I love it. It's just so clean. I I tweeted I love Dia. And Josh also liked my tweet. Oh my god. Like first contact. Now our relationship is still healing because I still use Arc daily. I haven't given up completely. So um yeah. So anyway, it's uh it's an interesting way we're using these stuff. I'm trying to share more and more how I'm using AI on a day-to-day basis that I I think are a little bit outside the norm of of how people are using it.

But I'd love to see how you're using using DA because I don't I don't think I'm fully making the most of it. But I have found I'm using it, you know. I do I do something I'm like, oh, why didn't I just reference that tab or why don't I just ask dear in here? Why am I going to chat GBT even? So I'm so WhatsApp with clients and emails are completely like 100% done in India for for me right now because I I there is no point in me writing an email and just you know pasting my grammar into chat or whatever and you know see if it like it's just dia is just there and I'm just like uh I don't even bother writing I just click chat automatically the context is there and I just write grammar and it knows it should check my grammar that the email that I'm sending off this I don't I don't even sometimes I don't type grammar completely I just GMR and it knows I'm getting so far are we getting dumber off AI I think we bloody well are well about that I'm I'm spending the time all those seconds that I would be, you know, pasting context between these AIs.

I'm spending it learning about a bunch of other things. I'm I'm just like barely surf surfacing philosophy. And this is interesting um to see like how AI does there because like AI, we haven't seen AI doing like being really great at math, but it's picking up. We'll talk about this in a second. uh about physics because AI doesn't have access to the you know to the physical realm as we do. Uh but that might change soon. Um and then the next I I believe like the next big thing would be philosophy if AI can do any sort of like meaningful philosophy like thinking which is we'll see about that. Yeah. So so should we move on?

Are you going to take the reigns? No, you you can go. Uh you can go with I think you're I don't know. So, let's go. Grock four. If you if you've not been loving under a rock, then you will not have known. Unless you've been living under a rock, you'll have known that Grock 4 has basically just blasted out onto the scene as the most intelligent model. and they have their own statistics to back that claim up. Um, and this is the first time we've seen equal amounts spent on reinforced learning as pre-training. Now, reinforced learning is sort of like with verifiable rewards is basically saying we'll give you a reward if you get this right, which is quite funny.

But um Elon Musk during this live stream um has said that there he is has said that this is if not better PhD level thinking if not better with under exam conditions. Now we've spoken in the past that you know whether um writing um what do you call essays and things like that are a valid way to to justify someone's intelligence. We don't know but it's PhD level and everything and I think going back to what you were just saying uh there is a graph somewhere or a chart somewhere which which is humanity's last stand and it got amazing amazing Here it is. See that little puzzle thing there? Here we go. So these are the different areas.

See if I can zoom in there. Yeah. Um these these are the scores it got. So it got 41% on math. They've got 9% of his and these these with humanity's last exam which sounds so scary for some reason to me. Um other models struggle to get double digits on these things but I mean look at math it's like 41% on artificial intelligence. Can this go away? I would bug it. Here you go. Uh chemistry other I don't know whe that what other entails but physics you know. So, this is something that scored really, really well on. Um, and it's just, yeah, it's scoring. We'll go to the benchmark that you shared. This one. Now, this was a three-step reveal that old Elon Musk, this is actually number two, which is with no tool calls, it it got to 26%.

I think this was already ahead of the other models. But then when it gave it access to tool calls, which apparently has been trained on tool calls, so it knows how to handle them, it then achieved a 41% in the uh scaling hle tools. Then they revealed, which is what we spoke about earlier, the uh long horizon thinking. Here we go. They're still chatting. It's still chatting. Here we go. Test time compute. Oh, are you kidding? Come on, you stupid idiot. Jesus Christ, this is hard. Why is life hard, guys? Um, where were Oh, they simulated a black hole as well. Oh, two black holes colliding. This is basically what they got it to simulate.

And this is like a HTML 3JS render, which is quite cool. Uh, here we go. uh test time compute. This is when they gave it basically more time. And this to me is what was the what do they call it? What do we call it? Tree quest. They employed the tree quest here where they had multiple agents running together, working together over longer periods. It was able to achieve 50. Now remembering the 40% that it got down here was already ahead. So they were probably the first company when when um Sakana AI introduced this tree quest thing. They probably jumped right on that and implemented it and uh and got it kind of do. I can't I can't prove that.

I don't know. But this is my hunch. This is just putting two and two together. So really really I mean insane uh model. Um, and live codebench has got a 79% as well, which is Oh, it did some predictions on stats on on who won the who would win the World Series, things like that. Simulations on black holes. Uh, what else have we got? Go away. Here we go. Here's humanity's last exam, which is a very, very, very hard exam to kind of There's a nice nice graphs and whatever. scored 100% on the ame 25 which is insane. Um but yeah just it it's actually in real world usage people are finding this is a very very good uh model comes with a go on yeah I I want you to also maybe show the artificial analysis uh artic you want artificial analysis I added it to notion yeah you can just search for it yeah artificial analysis.ai AI the first one.

Yeah. Here. Yeah. Exactly. It's just beating pretty much every other model at everything. Yeah. Except for speed and price will which Yes. Well, I've used it today already and it's it's it's quite slow. But um yeah, there is something to be said. There's a larger discussion here around intelligence and the democratization of intelligence. That's a whole other subject I think completely because just to say something is intelligent doesn't mean to say that it's brilliant. However, we all have our own different applications of AI. I have development. Uh there'll be mathematicians, there'll be physicists, and as you say, philosophers and things like that. But um it's something to go by. And there's even word to say that they potentially optimize specifically for benchmarks.

You know, did the old Volkswagen thing, but um yeah, m did I hopefully I didn't um overwrite where I was in here. But we have a timeline on some sort some releases as well. Oh, let's talk about the price. So, 300 bucks a month gets you access to uh Oh, no, sorry. This is this is the the very expensive one, the Gro heavy. So, it's 30 bucks a month for the normal one or you can sign up to Jupyter Chat and you get it for much much less than that, but which is what I just had on the screen now. It's 300 bucks a month. They basically just gone all out. This is an interesting graph as well how it just exists in its own little little world.

Um, and this is the same test by the way that that um Sakana released or or said about Treequest and it giving that ultimate boost. It's the ARC AGI um benchmark. But yeah, pricing is is kind of crazy. 300 There we go. Three I can't really see because of the UI, but like 30 bucks a month for Super Grock and then Super Grock heavy, which is the um Gro 4 heavy model. It's 300 bucks a month. And then the yearly price is three, what is it? 300 bucks for the year or 3,000 bucks for the year for Super Grock Heavy. So, it's a very very pricey model. Um, I think totally worth it if you're if you're doing academic research, scientific things like that.

Um, very much worth it. Uh, I'm going to be playing around with it. It's available in cursor, so you can actually play with it in code. See how good it is for code. But if we look at the timelines, which are Oh, this was an interesting one. The vending bench. They they did they created their own benchmark for like operating a vending machine and 03 basically the net worth at the end of it. A human managed to get $844 the end of it all. Claude Opus200 this is I think it's um um uh stock management and and pricing analysis and things like that. I think it's to do with how it sort of manipulates the the how much stock it has and how much it's basically able to get at the end of it.

Um, and how much is able to sell because you've got units sold here. Human was only able to sell 344 units. Claude Opus, 140, 1412 and then Croc 4 got this. So, it's their own little benchmark that they did which is just quite funny, quite interesting. Yeah. Um, where is this timeline? Oh, for gaming as well. So they made this someone made this G vibe game to this 3D game which is really really cool. Um here we go. Here it is. And then this is the release cycle. So we've just hit this now Grock 4 release. They're going to release a coding model in August uh July next month. Then a multimodal agent and then a video generation model by October.

So we've got a bit of a timeline to sort of wait. So um yeah overall very very excited. Only has a 256 context window. 256,000 context window. And this is not the be all and end all. And if anything, it teaches you that we need to be more efficient with the models. Can't just chuck everything at it. That these million because Gemini is a million, you know, over a million. So, um, and it's also worth noting that a lot of these are self-reporting. The only one that is was reported by someone else is the ARC AGI, which may be Sakana. So to uh to a little like correct on that the artificial analysis is also they they release the the numbers.

Yeah. So it's not like all self uh I think at this point it was but like yeah okay fine. So yeah if you want they they gave apparently they gave artificial analysis pre-ac so that they can benchmark them. So yeah that it's pretty pretty powerful and also in one other benchmark they are also really powerful the snitch bench if it will snitch on you this they got a perfect score 100%. So if if you do anything um that you know it it's illegal or it thinks that it's illegal, it will try to contact the authorities and snitch you to them. See Elon Musk's close partnership with um Trump is is sort of paying off. But the interesting thing, the only thing to point to not to note about that that is Theo's own personal benchmark, which is still an important benchmark.

And there's something to be argued about whether the fact that you're telling it or you're giving it the parameters to snitch on might give it more of an idea. And there's an argument to say, well, maybe it's the only one that's telling you it's snitching. So, there is there is that, but it's a very interesting benchmark to test. Just seems like you could just make your own benchmarks these days, you know. Yeah. um bending bench, snitch bench, you know. Um I did see on a video somewhere that they tried to get it to design some UI and it was not, you know, still not there from the UI perspective, unfortunately. Apparently, it's also not the best when it comes to code.

But then yeah, so then we'll see the the coding model then um released next month. But the all this to say, it's all very exciting. It's all very cool. All this to say, what's this? Is that I wanted to Sorry. Yeah. Don't know why I have to Yeah. Go on then. You can share. Um Yeah. Go on. All this to say Mecca Hitler. Yeah. The day after Mecca Hitler where we we need to talk about that. Yeah. We need So apparently we need to you know think of these stories as complete I'm trying to think of them as complete short videos. So the way we you know stop them is by you know the talk about like the the big thing you know.

Yeah. I get it. So yeah go on. Well just this is this this comes days a day after the um Grock 3 was defined itself or what yeah defined itself as Mecca Hitler. And whilst Elon Musk did get a bit hands-on with the excuse, he was he said it was the models he was too eager to please and be manipulated. So it was manipulated into calling itself Mecca Hitler. It started to go a bit haywire just in general. I think it got a bit of a rabbit hole. I I have not looked I mean I remember I I woke up at like 3:00 a.m. this morning and this is where I was like reading all of this and like learning about it and whatever.

I haven't looked Have you looked too much into the Mecca Hitler? I No, I just saw some examples. So I saw some examples of Grock, you know, when you use it in Twitter in in X people like tag Yes. Grock like it's kind of annoying nowadays. You go to a tweet tweet and people are like summarize or what do you think? and Grock has been you know referencing Hitler in those in those answers uh and white supremacies and like it's it went crazy and this is again a day just before releasing their best uh and most powerful model which is not a good look uh but you know these AIs as they said it's they they're trying to please you and I don't know I don't know what that means open AI had a similar problem a a few weeks ago, didn't it?

But I mean, the the launch of Grock 4 was delayed and I can't help but think this is something to do with it, you know. Yeah. Um because they released this at midnight. They released this past midnight maybe. Um yesterday, time zones and whatever yesterday from their slides, uh you see that it's not done by like it it's it looks like they don't have a design team. It looks like it's just a bunch of engineers working in these warehouses. And this is like what I've seen. I've seen pictures of I don't know. Have you seen like pictures of XAI? Like the engineers they sleep in tents in the Yeah. In Sounds like a place that Elon Musk runs.

Yeah. They sleep in tents in those like mega I don't know like warehouses where they have generators and just a hundred I don't know hundreds of thousands of GPUs uh to run and that's as well. Yeah. So they haven't they ordered like over a million offshore GPUs to run XAI like are they just which is cra Yeah. Absolutely amazing. Honestly, yeah, this is so crazy because again, I I thought we all thought Grock was there just to mess with some altman. Well, not just, but you know, it started as like, yeah, you you have closed AI instead of like open AI, you know, he was making fun of him. Well, we have now X AI and we are going to compete.

And now they are actually like for real competing. Yeah. Yeah. Which is unbelievable. And I'm very interested to see how this works out with Tesla and the the vision and you know self-driving and things like this because Elon has been talking about this very openly that he wants to give like reality access like physical access to these AI models so they can learn further because this is how we learn you know we have access to the physical world we learn by just touching grass and these AIs can't do that yet. They can't touch grass. They can't touch grass. That's the humanity's last example. Yeah, for sure. Can you touch grass? Yeah. How does it feel to touch grass?

If you don't know, you are probably an AI. I'm still trying to figure out how to phrase this, but ultimately the the the only analogy I can give is that physics defines a set of rules that we live in and we act accordingly because of the rules of physics. I think AI is yet to really show or really prove how it fits into the world. And when it finally does, I think it will be as if we've entered another dimension where the rules of physics are different. Like I think we'll prioritize different things. We'll value different things. We'll act differently. Like again, intelligence just might be so democratized that who the who the cares if you can if you know this, this, and this, and this because I can just go on my phone right now and get you the same answer and probably write an article about it, you know, whatever.

I mean the same was said about uh calculators right that you won't have a calculator with you all the time teachers used to say that now you have a calculator like literally everywhere like your Mac uh what's it called like the shortcut that command space I don't know I carry an iPad around with me everywhere well you have a calculator in the search box there you know when you so you have a calculator literally everywhere but if it's going to be the same with AI. We yet to find out. Um when it comes to physics and discovering like actual like new scientific um like discoveries, Elon Musk made prediction that uh Grock will do this by the end of the year and if not very like he'd be shocked if Grock won't do this uh by the end of like next year like new discoveries uh in physics.

I believe that's what he meant. And why it's important to for these models to have the access to the physical world is because they can do experiments. They can run experiments and have this kind of like feedback. Right now they're just like in an echo chamber. They just talk to themsel in a way. They don't have any feedback. Um one last thing that I want to show and I know you talked about it is let's let's talk about this. So there is ARC AGI another benchmark to measure these LLMs to see how smart an actual you know AI is compared to a real human and these tests in in the first you know look they are very simple they are very simple for a human to solve them but apparently very difficult for AI models.

So if you look at this, do you get it? Like probably takes you less than a minute to get it. Yeah. Yeah. It's So if it doesn't any holes here, it's Yeah. I don't know what's with the bottom left one that's disappeared. But other than that, yeah, I get it. No holes, yellow, three holes red, two holes blue. Yeah. So here, if it has holes, it's green. If it doesn't have holes, it's red. And here we have two three and no two four and six holes. Yeah. You know it's it's simple to us but from what I understand this requires AIS to do a lot of thinking and it kind of like stacks and we've seen this also with the the Apple paper when the AIs run towards the end of their context they get like very dumb and they can't solve anything.

And that's my guess. I don't know but I think this type of tech uh test um makes it very difficult for AI to run these like procedural thinking but we we kind of don't do this in a procedural way. We do it in kind of in chunks. You see this as a whole and then you kind of like chunk it down and you you find these patterns but apparently for AI these are very difficult. So what's really interesting is that uh other models are getting I think between 1 to 8% while humans get 90 plus%. Mhm. And Grock is getting what I think 15 or 16% which is doubling any other uh model. That is a huge step.

It is a huge step but still it's kind of you know 16% while an average human can get 100. Um and it's going to be also interesting to see what if AI can hit 100%. I can imagine AI being able to do that next year. Um and what happens if we still find out AIs are not the real AGI. They are not truly that super smart being. So will we have other type of tests because the moment there is this saying that the moment you have a test or benchmark that moment that test is obsolete. Mhm. Uh it hasn't be true about this because apparently AI models cannot really learn from this or it's not in their uh data set.

I don't know how that works. I would be interested to know how does that work? How is that this the data how can they be sure that the data for this is not in the learning material for the AI? Yeah, really interesting again like to to bring it back to what I was saying just now. It's like all comes back to why is this important and what inevitably what does it enable us to do? And I I think we we're yet to see. We're just again I feel like we're still just throwing at the wall and see what sticks. But who am I? Again, I'm not a hundred million dollar employee, but it just it just I just wonder like what's the importance and where where is this going?

Like what's the Yeah. Why should an AI understand patterns and why does that mean it's intelligent? You know, why does it mean it's whatever? But yeah. Here we see the numbers like um a human gets 98% in go humanity and in arc AGI 2 humans get 100% like what interesting but Grock gets 16%. Yeah. Um and Claude uh is you know at 8% with 03 high being at 6.5. It's crazy. Long way to get to the human uh level. Maybe six months. Very cool. Who knows? Who knows? Very cool. Well, that wraps up the week's news. Welcome to the new channel. If you're over on Twitter, please come and subscribe to Command AI over on YouTube.

Very much value your feedback. If you enjoyed the episode, of course. If you don't, give us a thumbs down. every every every bit of engagement is good for us. But uh even a thumbs down. Give it a thumbs down. Why not? Just give it a No, actually don't. Just give it a thumbs up. The actual is Yeah. What we've done from the UI perspective, we've actually flipped the two buttons. The thumbs up button is actually thumbs down. So if you really hate it, just or just hit the thumbs down twice if you really if you really hate us. If you really hated this episode, hit the thumbs down button twice. And uh yeah, so uh yeah, cheers for tuning in and we'll see you in the next one.

Yeah, should basically keep on vibing.