—
TIMESTAMPS:
0:00 Preamble
06:24 show starts / Intro
09:03 Codex in ChatGPT app
15:02 Osaurus - privacy first local AI app
26:03 Notion goes all in on AI agents platform
30:04 Cursor's Composer 2.5
41:25 Karpathy joins Anthropic
48:35 Google I/O announcements
50:03 "Neural Expressive" design language
51:15 Gemini Omni
56:51 Daily briefing
01:00:08 Gemini Spark
01:04:28 Gemini Mac app
01:06:39 Antigravity 2.0
01:14:25 Gemin 3.5 Flash
01:19:50 OpenClaw x Grok
01:24:14 Perplexity query-aware context compression model
01:29:17 Figma Design Agents
—
Unlock the full potential of your online presence with Kabarza and Samuel—experts in web design and development (respectively), powered by cutting-edge AI solutions. We blend creative design with advanced tech to deliver smart, high-impact websites that stand out. Ready to elevate your business? Contact us today and see what AI-driven innovation can do for you!
LINKS & RESOURCES:
Website: https://cmdaishow.com
Check out Kabarza's amazing work: https://kabarza.com
Visit Samuel's website for more: https://samuelgregory.co.uk
📷 Follow on Instagram: https://www.instagram.com/cmdaishow
—
HASHTAGS:
#ai #podcast #aidesign #aidevelopment #vibecoding #webdesign #webdevelopment #ainews #webnews #designnews #devnews
Transcript
Oh god, didn't see you there. And just joshing. Just joshing. Check everything's going on correctly and as expected. Where are we live? We are good to go. Good afternoon, good morning, good evening. I don't know about you or where you are in the world and depending on where you are in the world, this might not be actually that exciting, but it is so so warm and sunny here in London. And I'm very excited about that. And it's this thing called bank holiday weekend, which again, I don't know where you're from. Let us know. We have these random holidays where it's like, yeah, it's called a bank holiday. So we can get a day off and there's loads in this kind of like first or beginning of the second quarter of the year. I don't I don't know where we are even. So yeah, this weekend is a bank holiday. It's a ve going to be a very very warm bank holiday. All of Britain is loving it right now. So I'm going away. I'll be getting a train at the end of this session. and uh heading into the the countryside. And as I see some of you start to trickle in, let us know where you are from. But also, Gabbaza is not here this week. He will be he's attending to some things. So, he's uh skipping this. I think he's skipping next week as well. So you you're stuck with me this week and next week. So what that normally means is we don't yap. We just get straight on with it. But you don't get that color. Missed the nice weather in London. [snorts] Yeah. Although it was nice yesterday. I think I remember I was in I was in my my room all day working away and I went went downstairs and I was like whoa is the heating been left on and uh no it was just it was just that that thing that's in the sky that provides us warmth and nourishment. So yeah, it was nice yesterday down in sunny uh sunny south, but you know I was on a plane all day yesterday. Yeah, that's the that was the day. What were you doing here? What was here? What was going on? Um I saw like there was a I think there was a Claude thing going on, I think. Or they're they're announcing something that's going to be in London. I don't know. But um yeah, let us know what you're doing. Yeah. Okay. Yeah, you're right. All right. That's sick. Um how like what did you do? Was it like an announcement or something? cuz I was really confused by the um the uh just the post that I saw and I was like is this is this going to be happening or is it happening? Like it seemed like a teaser but I couldn't see oh and and were you invited or did you need to sign up to be invited? And finally, where did you find from? Name like Skylar Skylar Kitchen. I don't know that sounds kind of Scottish to be honest, but let us know where you're from. was sick like Web Flow Conf. Nice. Any any sneaky announcements they made of future upcoming uh you know products. [snorts] We'll uh we'll be getting on with this in just a second. Let me just sort some stuff out. Whoa. San Fran. [ __ ] hell. That's quite a way to come for CL. Were they doing it in San Fran? like what was the big what was the big um claws and deep dives? Yeah. Yeah. Just like a keynote thing then I guess with some breakout rooms I think. Yeah. That's a long way to come for that. Sorry you didn't get the sun. [snorts] Hey, and you probably get some in San Fran anyway, right? All the time. Cool. I think uh I think we're going to roll with it now. So this week, North Carolina and got noticed. Nice. Not [ __ ] my life. You've been traveling, so that's all right. That's That's a good life. That's a good life. Cool. So, um going to maybe skip that one. We'll breeze over it. Codeex goes mobile. Open Eye brings its coding agent to your phone, taking um aim directly at Claude Code. Osiris or Osaurus, sorry, Osiris. Osaurus. Osaurus, a new Mac app that lets you run local LMS and cloud AI actually. Uh side by side, uh focusing on privacy first and Mac first, MLX for you local AI dude and dudetss out there. Notion becomes an agent hub. It's no longer just a note-taking app. It's now an orchestration layer for AI agents. Co cursor composer 2.5. A smarter coding agent. Uh, that's 25 times uh trained on 25 more times. [ __ ] that one up, didn't I? Composer 2.5. A smarter coding agent trained on 25 times more synthetic tasks with some wild reward hacking stories. Um, does look pretty good actually. And I don't know, it's interesting to know anyone's used it. Kapathy joins or karpathy. Sorry, I keep getting this from Carpathy joins Android. Uh Jesus Christ. Carpathy joins Anthropic. One of the most respected names AI just picked aside. We'll dig into that one. I got quite a few thoughts on that one. Everything Google IO, you've got Gemini, you've got um Gem Gemini Goentic, uh 3.5 Flash, loads of little tools. They've um what was it called? Omni, they've also released like an open claw competitor. So, really cool there. Um, [snorts] Grock is making waves in the industry now and in a big way given it's probably supporting half of their infrastructure. But, uh, you can now natively it now natively supports openclaw or openclaw natively supports gro 3.5 perlexity compression uh, model. I we're going to read through that one. I've not looked deeply into it, but um and then finally, Figma's AI agent, which looks really interesting and has uh maybe answered a few of the thoughts I've been having recently about design and AI. So, all that and more and this week's Command AI. I'm Sam. That's all I've got for you. [laughter] No, no Kabaza this week just because yeah, he's got some stuff to attend to and uh unfortunately he won't be joining us this week or next. He's on a I know he's on a holiday next week, so that's pretty cool. So, let me get some screen sharing going on. First up is [clears throat] uh maybe I delet maybe I remove Oh, no, no, we've got it. We've got it. Um I was like reading about it and I wasn't too interested in it. So, Open AI brings codeex to your phone, which is they they've made some really interesting strides in terms of like control through codecs. Uh, recently I've seen some updates with regards to computer use and stuff like that. Doing it a lot better seemingly than what Anthropic does. But, Codeex is going mobile. The coding tool which OpenAI launched approximately a year ago has now been integrated directly into the chat GPT app allowing users to monitor and manage the development workflows remotely. Now from what I understand this isn't just being able to monitor like uh remote coding agents. It's actually monitoring the agents that are running on your machine. So it's like remote control which is what uh Claude do um but just seems better integrated. I'm not a chat GBT subscriber, unfortunately. I don't have I, you know, I I I mean, I looked into it. I just couldn't see uh chat GBT. I literally couldn't see any way to to do it. It's available on all plans, obviously, except the the free plan. I should probably um take a look at their super cheap, their go plan, and see if there's that includes codecs. But yeah, I don't know. And the new function allows users to see their codeex live environment in any devices where it's running. So whether that is on the cloud or it's on your machine, you'll be able to see them all. And then not only that, but you'll be able to respond to the messages and and do things with it. It's not just like dispatching commands or setting up to go do something and then exit. You're actually able to work with it, which is quite nice really. I think that they're really making strides here. Here we go. philanthropic released the similar feature remote control which allows users to remotely monitor Claude's code's work from afar. I mean if you're a fan of uh claw code then it's obviously a really cool feature. Um I don't know let us know are you are you coding on your phone? Are you setting tasks off? Because I recently discovered I didn't really really discover it. was enlightened by the usefulness of just tagging Claude in like a GitHub issue, like creating a GitHub issue and then having that become the prompt essentially tagging Claude. Sometimes that just needs the title to be fair, but you obviously got a description where you can provide more context. Um, but I got to be really on this bloody uh this switching around of this thing. But yeah, um it was quite nice just living in GitHub and then you've got like a you you can um work with the say the um the native issues feature of GitHub obviously pull requests and things like that like it works quite nice like that. I quite liked it. But how many people are actually coding from their phone or maybe you've just popped to the shop or you pop to the toilet and then you're just seeing what's going on and then you you know whatever. But I see people uh walking around with their laptops open and stuff like that which Jules actually Jules took shots of that. Um Jules uh tweet laptop. Let's see if see if this brings up anything. First post is up. No. Jules Jules agent. Let's see if it's on Twitter. Um, people are like walking around with their laptops open just so they can keep their agents running and stuff like that, which is just ridiculous. I mean, come on, guys. Jesus Christ. Go touch some grass, enjoy the air. I mean, I'm not the best proponent of doing that, but I'm not walking around my laptop open. Mainly because it's too freaking heavy, but you know, that's is what it is. Yeah, I can't see it. But regardless, codeex on the phone is is definitely uh definitely welcome. um having everything run from Slack is the future, you know, and and that that was one of the things that um that was one of the things that like made me think about it like in a different way like chat like treating it like a coworker I think is what what the this the common thread there is. You're just you're just having you're treating it like a co-orker that just picks stuff up and does stuff when you need it to. Um, which I I think those who use it, like I mean you'll be able to you'll be able to uh attest to this, but those people who use it like that will will swear by it. But until you've got to that point, I suppose if you're like a solo dev and you don't work with a team, you're probably not used to like living in Slack. For instance, I used to live in Slack every single day when I worked for a company since working for myself and working solo. never touch Slack. Um it always annoys me when uh I work with San Francisco uh clients and then they want to add me, but then I've got to upgrade my Slack literally just for that one project, you know. Um but yeah, I agree. It's just having it having it where you are as opposed to like going to it. I think again is the common thread there. But um yeah, again, maybe that's your phone, but I you'll see like my phone is like black and white. I try and use my phone as little as possible during work hours. I'm not sure how effective that is to be honest, but um because then I end up just answering WhatsApp messages on my iPad or like using Telegram on my iPad or something like that. But point is I do try and stay away from my phone um during work. So yeah, not really for me. But codeex on the app might be for you. So yeah, I welcome it. Um, next up, this one I'm quite excited about. And I've got to set up my Mac for this. So, I've been getting big. Maybe you're here because of the the fact that I've been getting big into local LMS on my own my own um channel. And the big thing in that is GGUF versus MLX. MLX running local models being like the native way to do it. Um, what's going on here? No, that's the wrong screen. But either way, um, get it all up already. There's a new local agent app in town, um, that is set to take on Olama. It's called Osaurus, and it's a really cute dinosaur themed app that will not only run local LLMs, but also cloud LLMs, putting privacy first. So, let me clean up my dashboard. I will show you the ropes. And let's have a look here. Jesus Christ, I've got lots of windows open. Um, [snorts] let's put let's put my code in the background there. Uh, [clears throat] I've lost you. window. Let's go. Oh, entire screen. That's what we want. And we're going to go here. And we're going to go here. And this is minimize you there. This is Osaurus. and it's really really cute. So, basically, you set up these little dinosaurs as your agents. Let's call them harnesses. Like, right now, I've got a co I've got a coder agent set up. If we go to settings here, [snorts] um you can give it a a system prompt. You can set it a specific model. Uh all this that and the other, whether you're doing chat, whether you're doing code, whether you're doing, you know, whatever. and it pipes into local LLMs that are MLX focused. Now, if you don't know anything about local LLMs, MLX is the Apple specific uh quantization. I don't even know to be honest. I probably should know a bit more than that, but um they run much much better on Macs. And I've done these tests. I literally put GGUF which runs better on the kind of Windows I guess or the um non non-metal GPUs uh put them up against each other even with this new um cache uh turbo quant the caching thing even with that like it's just MLX just wins hand down. So if you're running any local LLMs, you you make sure you do MLX and oh yeah, MLX. And all this is doing is pulling down from um Hugging Face anyway, which is a repository of like uh AIS basically models. So I pulled down this one, the Neotron 3 Nano Omni 30 billion parameter model with three billion active parameters. That's how you break that down. So I downloaded this one. I've got 64 gig of RAM, which might sound like a lot, but I've got an M1 Max, which is, you know, amazing for other stuff, but for AI is dog water. It's rubbish. And it's Yeah, it's cool that I can run it. And ultimately, if we come out here, we're on our coding thing. We're on our coding agent. We select uh the Neotron and we can go. Hey. And that's going to take forever now because and my fans are probably going to spin up whilst it's doing it. I'm going to get toasty and warm especially because I've left thinking on. But yeah, and but you can also you might have seen as well while it's doing that. Um let's go settings. You can also pro put providers. So you can link it to anthropic, Azure, Deepseek, um Olama even open router. So it's and it's all privacy focused. And I think that what that will mean is is that when you're using OALMA or you're using Open Routter, there's a special flag that you can pass it in your request that tells it to use a gosh, trying to break it down if if if you're new to all this. Um, these models run on different infrastructure. Like, uh, Claude doesn't always run on Claude's infrastructure. Well, actually, they've just struck a deal with SpaceX. So, it's running on SpaceX's. Sometimes it's running on Azure. Sometimes it's running on bedrock, which is Amazon's infrastructure. Only some of those infrastructure allow don't collect any data. And you have to pass it the flag to like you have to pass open route or a flag to say don't send me to any providers that are privacy uh that are are collecting my data. Um long boring way to say that because this is privacy first I wouldn't be surprised if the uh if that was the case. So really cool there. I'm going to make sure I bring you up down here in case there are any comments so bear with. Um, and what I really like is what I've been doing is it I mean, let's just go through actually. You've got the agents, you've got plugins as well. I've not really Let's have a little browse. Apple Music. So, it can tie into all of your Apple stuff, which is really nice. Apple Maps. Look, that's really cool. I haven't looked into these. Again, using the local LLM as the sort of interpreter layer. Um, sandbox. really well featured tools, uh, skills. This is cool. Um, and I'm I'm guessing if you've got an agent, you can then Here we go. capabilities. You can give it certain skills. So, that's quite nice as well. Really, really nice actually. This is so much better than Lama. Um, memory even. This is really cool. Suppose Lama's nice just cuz it's so simple. But, um, the one thing I'm really interested in is this one. the server. Um, and what this is doing, this is running on server 1271337, and I've generated an an API key. And in the background here, I've got Pi uh, running. And I really want to do a video on Pi. Let us know if you're interested in Pi. Um, really cool coding agent. Really, really minimal, but I've set it up to run on this endpoint that's running. Um, Osaurus. I can't see that because I've got a little popup which you probably can't see, but I'm hope actually. Let's just go boom. Yeah. So, I'm running the Neotron out. Um, Probably should have cleared the context because it's got all this rubbish in there. But this is pi. Um, and hopefully this is now working away. Maybe we can get up the activity monitor and GPU history. I mean, my my memory pressure is pretty darn high. There it is. It's ramping up. So, this is now using a local LLM on Pi. So, this is completely private, not being sent to anyone. Um, you know, telemetry is one of those things that people can kind of sneak in. But from what I understand is that, you know, again, if something's privacy focused, I sure as hell hope they're not doing it. But this is all working locally. again, slow as balls because I am running an M1. But Osaurus, this is I'll be definitely making a video on this because I think it's a really nice competitor cuz Olama, see, it's even still thinking. Look, Olama are um they've only released one model under MLX. The rest of them are GGUF. So to get MLX, you have to go to LM Studio. That's becoming a bit bloated and a bit rubbish to be honest. And um I moved over to OMLX. But that's it's quite nice cuz it's so so simple. I wonder what's with these O's. Um it runs in the browser as well. So, if we click up here, go uh start server. Just I'm hoping I'm not doxing myself anywhere here, but heyo, whatever. Admin panel. So, now this is now running. Um, what models do I have? So, I've got all I've got these models going. Uh, how do I even run these models? I've forgotten. Um, well, I guess this is now running on 8,000 port 8,000. So, if I want to get it running in clawed code, I just do that. Uh, would not recommend running local models in clawed code because claw code just has so much context before you've even hit before you've even hit return. It's got so much context. So, yeah, OMLX is quite a nice thing. But, um, this is literally just local models. So, if you do find yourself darting between local models and your uh like remote models, Osaurus looks like your jam. And it's super cute and really nice to use. So, there we go. He's responded. Pi, is he working? Yeah, this is this is the power of M1 Max, by the way. So, yeah, really, really like it. And I'll be probably making a video on my channel about Osaurus. So, pretty cool. Um, what have we got next? If you've just joined us, Kabaza is unfortunately not going to be with us today. Uh he's attending to some stuff. So, you've got me, which does mean that we can go through things a bit bit quicker and a bit more succinct because we're not waffling away. Um, what do I want to do? I want to go to notion. So, yeah, pretty good. You obviously need a good machine and I've got tons of videos on my computer, you know, setting up um Macs and things like that. Uh what what you might need and the sort of strengths computer you might need. So yeah, go check that out. But yeah, this one's a really interesting one. So, Notion have just launched a developer platform letting teams connect AI agents and external resources together, custom code directly into the notion workspace. It's pushing deeper into the agentic productivity uh software. So, I run so much of my company runs on notion. It's not the most ideal, but it's super bloody flexible and I can plug into it and I have integrations and this and that. It's very flexible. So what notion have done let's um share my correct window now. Mute this bad boy. Wonder if there's anything to show. I don't think there is actually, but you can basically set up this can go away now. Uh you can basically set up integrations. I think that's what it is. Integrations with all of these softwares and they work directly with the um database. Maybe we can show you that the databases that you have set up. So, they've even look, they've even um got a little dashboard here that can just link everything up, pull data down. There's a Stripe integration here somewhere. Um the the agent orchestration's really nice as well. So, you can set up all of these different providers to actually do the work itself. You can schedule them. It's it's quite I mean you know I'd have to dig in a lot more to to properly understand it but they call these things workers as well which again just go off and do stuff such as pull stripe data or you know and and then sync it to the database. So really nice um because there's a lot of complaints about notion neglecting their developer audience. They used to be very developer friendly and now they kind of went down the, you know, conglomerate route, but now they've kind of pulled it back a little bit, starting to support it with the CLI, with agents and stuff like that. Um, and it looks really interesting. My only issue is, you know, because they're so late to this game of, you know, setting up allowing you to set up agents and stuff, people have already found other ways to do it. Like I've got loads of nan workflows that just read from my database, push into my database like um obviously clawed connectors, you know, connects into notion and stuff like that. There's no re like I feel like it's a little too late for notion here. I'm welcome to be proven wrong and shown that actually there's some really unique use cases that you can only achieve in the notion. um agent that whatever you call it like the forgotten what they've called it um I don't know workspace hub whatever you want to call it most of this stuff seems to be able to uh be achieved elsewhere however obviously people are pining for it I quite like the CLI I think that's quite that's quite a nice thing if we try and find that you can like set up agents from basically your local machine as far as I See uh and then you know manipulate data create create create create workers even from the CLI and stuff like that. This is quite nice. Uh but that the other stuff synchronizing data I don't know little little too late there notion I think to be honest little little too late. That's all I've got to say on it. Let us know if you can think of some cool use cases for it. But uh yeah, Composer 2.5. This was a pretty big deal actually. Uh to be honest, it's just Kimmy K 2.5 or yeah, Kim K 2.5 under the hood. They've just uh done some reinforced learning to uh to get it better. Um let's share this tab. Let's go through the article and see what they are saying. Introducing Composer 2.5. Composer 2.5 is now available on cursor. Who is a cursor boy? Because I have not touched cursor in ages. I've gone back to VS Code and even then, yeah, I've gone back to VS Code. I still have it open on all the projects that I'm running. Um, but yeah, it's a really dire space for them to be in because I think they've turned their tool into a bloated mess, but whatever. Let us know if you're a you're a cursor boy. Uh, substantial improvement in intelligence and behavior over cursor 2. I should damn well hope so. Is better at sustained work on longunning task. follows complex instructions more reliably and is more pleasant to collaborate with. Interestingly, they haven't mentioned the speed because that was always the thing with Composer 2 that it was just crazy fast for its intelligence. Crazy fast. Let's look at the benchmark. So, Terminal Bench 6 point uh 69.3 be wow just shy of Opus, but to be honest, Opus is not that great for Terminal Bench as can be seen with GBT 5.5 here scoring an 82.7. So, you know, not not a great show of prowess there, but a welcome improvement over Composer 2 SWBench, which is what I'm interested in. I'm not too interested in benchmarks, but they they damn well give you a good idea if if I can think. Well, Opus 4.7 is pretty good at coding. It's getting really close to Opus 4.7, which is insane. I think you can only access this is in cursor as well. Let me try. Let me have a look at open ro. This might bring you along for the ride as well. uh compose I suppose this is a yeah it's not coming up is it yeah it's not coming up so this is a reason why you would use um a really good reason why you would use cursor to be honest just shy of um opus and I bet the price is very very cheap and then cursor bench which is really interesting because it's their own benchmark marks, but they're not really they're doing they're doing the worst. 63.2 versus 64.8 on max mode mind really drops there on extra high which I think is the one down from max and then 64.3 is X high 59.2 with GT5.5 but significant improvement over composer 2 as you would expect. We improve composer by scaling, training, generating more complex reinforced learning environments and introducing new learning methods. In addition to training composer 2.5 on more difficult task, we improve behavioral aspects of the model like communication style and effort collaboration or calibration. Sorry. These dimensions are not well captured by existing benchmarks, but we find that they matter for real world usefulness. Trust me, bro. This is quite a nice uh thing. So basically we're we're over I mean it's just this you know it's just this represented on the graph it's over um uh opus 4.7 in what are we talking about here like score um and then the cost is way down nearly zero. So what's that? That's between That's probably about 25 cents, right? Um, for this task, not even per like request or anything like that. This is the bit that I obviously highlighted. Composer 2.5 is built on the same open source checkpoint as Composer 2. So, it's still using Kim K 2.5, which is really interesting because I think we're in Kim K 2.6 now. bring you along for the ride. Um, I'm not signed up, but maybe we are to see K. Yeah, here we go. K 2.6. So, interesting. My gut feel in all of this point thing is that you can only realistically bump the major version if you've done some fresh training. So 4.7 um Opus 4.7 is trained on the same um checkpoint as 4.6 as 4.5. They've just worked out ways to train or reinforce learning and and post-training in such a way where they're able to achieve better or different results. So yeah, still using Kimmy K 2.5. Oh, look at this. Now I've been going through all this article and not even sharing my bloody screen. This is what what it takes doing on its own. So luckily I was quite verbal with this but this is what this is the this is the uh gr table I was talking about terminal bench 69.3 just under 4 uh opus 4.7 massively under GPT 5.5 but on coding bench 79.8 eight just under Opus 4.7. I think those are the key takeaways on that graph. And then you can see that demonstrated here. So, Composer 2.5 scoring better than 4.7 on this. Oh, it's a cursor bench 3.1 score. That's weird. I suppose it's Yeah, it is a little bit better. It looks a bit more dramatic here. So, a little bit better than Opus, but hugely cheaper. And then we got to here. Sorry about that. Um, composer 2.5 is built on the same open-source checkpoint as composer 2. Uh, so that's where we are now. Uh, 85% of compute for composer 2.5 comes from additional training and reinforced learning. So here, okay, so let's think about what I just said, which is the points are basically the same. Kimk 2.5 is trained on the same training data as Kim K2. So this is a training data of training data or reinforced learning off of already existing training data. At the end of the day, if they're getting this the results that they they're claiming to get, then who cares? But that's just interesting. Together with Space X AI. Wow, I didn't realize they changed the name. What's Where does this link to? Okay, I think that's a maybe a bit of a typo. I don't know. Um, but it is Curs's blog. We're we're training a significantly larger model from scratch using 10 times more total compute with Colossus 2's million H100 equivalent. So, I wonder what that means. What what what um graphics cards they're using if they're equivalent. Maybe they're cheaper, whatever. Maybe they're trying to keep it secret, but they've discovered um you know, oh no, they've discovered a new way to do it. I don't know, but interesting nonetheless. And then this one was found what I found interesting. Again, this is not where my specialty lies. However, I did think it stood out as something that you guys might be interested in. We train composer 2.5 with targeted cont uh targeted textual feedback. The idea is to provide feedback directly at the point in the trajectory where the model could have behaved better. So, however they're doing that, you could probably read this and tell me much more than I know about it. But ideally, I guess they're just saying if it does something wrong, they can just be a bit more specific about where it went wrong and then tweak things and change things um at the point rather than at the end when it has a bunch of extra context or something like that. So interesting and I wonder why other people wouldn't do that. It seems pretty straightfor like not straightforward but it seems like a pretty um you know yeah you should do that unless there are other downsides. Who knows? Now what I haven't seen is any mention of like speed. Here we go. Here's price anyway. 50 cents per million and 2.250. So, I don't know. There's a faster variant with the same intelligence for uh what's that? I mean, three, four times the input. I don't know. I'm not very good at maths. Regardless, if you need to get stuff done quick, they got a faster variant. I want to know how much how fast the base one is to be honest. I don't I'm opening cursor because I do have it downloaded even though I don't use it. I want to see if I get access to it under the free tier. [snorts] Um, where are we? Agent, I've got to log in. Is there a free tier? Is there a free tier of cursor? Okay, I'm in. I can select I can select composer 2.5. I guess auto mode is maybe composer or probably composer. Why wouldn't they use their own thing? I can select composer 2.5 fast. Sorry, I can't share it just cuz I'm um it's just a pain in the ass to keep switching between full screen and and specific tabs. But yeah, I can select I can select um 2.5 fast. I just don't know what my credit system is or anything like that. So, you might already have access to it. Let us know how fast the regular version is because you got a pretty fast model in the previous version. So, yeah. Um comp. Oh, there we go. That's quite a nice little uh addition there. Composer 2.5 includes double usage for the first week. So that's this week. So get in on that. Get in on it. Get in. Uh I'm going to check my notes here as well. Synthetic feature deletion task model must reimplement features that were removed while keeping test passing. That's just some of the training stuff I'm assuming. Yeah. No, pretty good. Pretty good. This is the interesting one. I'm not It's a shame. Um, it's a shame Kabaza isn't here because it's a chatty one. Andre carp Andre Carpathy joins Anthropic pre-training team after OpenAI co-founder stint. So, OpenAI co-founder and former Tesla AI chief Andre Carpathy has joined Anthropic to work on pre-training. He'll be uh build a team focused on using claude to accelerate LLM research signaling anthropics bet that AI assisted research is the path to stay ahead of open AI and Google. Now this this means more than you probably assume. Like this is the guy that's I mean he he's very he's such an intelligent dude and he f I mean look he worked he's founded open AI and worked at Tesla. He's the the dude to have. I think it's not it's not only a combination of how intelligent he is but how um able he is to communicate uh ideas and things like that whilst being so intelligent. I find him quite hard to understand actually. I feel like he's maybe a bit bit too smart, but there you go. That's just me being a dumbass. Um, but the fact that he has joined Anthropic means so much. He's basically saying that I agree with what Anthropic is saying. I agree with the what do they call it? the doctrine the the um the I don't know it's an American term but the the the foundations to the beliefs of the model it's he's saying that Claude have got it right by joining them and I mean I don't know what the what how that's likely going to change um what you know how it's going to shape the actual company inside obviously we don't know anything about that yet. But the fact that the the the godfather of AI probably at this present moment um is joining AI is a massive massive deal. Um I think these are so funny. There's a few of them actually where these like I mean look even here like The sports references are hilarious. I don't think he's It's not fair to say he's transferred though. Like he wasn't working for Open AI and now he's gone to Anthropica. I think he's been working solo for a little bit. I mean, correct me down in the comments. But yeah, uh I wonder also how much he's getting paid. I bet he's getting he doesn't need the money. Let's be clear about that. But um but it would have taken a lot of money for him to to go or maybe it's not a lot of money and he's literally like no he needs to get in on this this company's uh vision what they believe. Boris chiming in there. Boris churning in there. Uh there are easy ways to get mythos access for vibe coding. My guy, he's the guy who actually yeah coined the term vibe coding. Even though we were kind of already doing it, he was like, "Yeah, just giving into the vibes and just coding without even thinking about it." So love the term or not, it was literally coined by the dude who Oh, yeah. the godfather of AI. So I don't know. I don't know. Um I don't have much more to say on it. I just it's very very exciting times. It's it it gives people reassurance that Claude and Anthropic are the people um you know are the people to to get behind or at least their belief is is somewhat aligned with someone who knows AI very very deeply. Um yeah don't know let us know let us know what you think. Uh I'm going to read my notes here just in case I've forgotten something. Uh the next few years of the frontier be especially formative. Excited to be back to R&D. So that's just him what he's what he's saying. [snorts] He's going to be working under Nick Joseph. The most comput inensive and expensive phase of building the frontier models. Oh, I wonder if the models will get a lot better. Do you know what I mean? Like you know again we just spoke about uh post-training of communication 2.5 to get composer 2.5 you know is is you can do a lot with a model that's already been trained um or getting the most out of a model that's already been trained. Um okay so he's at Eureka Labs. That was the that was the last place he was at. What do Eureka Labs do? Sounds like a [ __ ] company name I have to say. Actually, we'll do it here so we don't have to Eureka Labs. We Oh my god, look at I mean, surely not. Surely not. This is This is a the website. Guys, come and give me a shout. Give me a shout. Well, anyone else? Yeah, we'll leave it there. We'll leave it there. Oh, actually, uh, Antropic also hired Chris Rolf for Frontier Red Team, a 20 year cyber security veteran from Yahoo. Um, I mean, Antropic are going all in now, aren't they? They are really going all in. I think uh um Dario has kind of just said, "Look, you know, we are we've got we've got the best agent. We've got the best a um AI. We now need to start investing." Which I wonder something I've been thinking about. You know, a few months ago, we reported on the fact that Anthropic was set to reach uh um uh profit or at least Yeah. re reach profitability before years before open AI. Um I wonder you know that came about because Dario has been a bit of a cheapkate kind of hoarding stuff or not spending but now he's just going on a spending spree. And the most interesting thing is is that nobody's leaving Anthropic either. None of the they it doesn't seem tumultuous. Now Theo has a lot to say about their poisonous dev, you know. um uh what do you call it um can't think of the word but the way they behave the devs and what they think of devs but the truth is no one is leaving you know they're all getting paid obviously well they're obviously you believe in the culture and culture is a massive thing in tech nobody's leaving so the the proof is in the pudding as they say the proof is in the pudding right we're going all right we're going all right I thought this be over in an hour, but we've managed to we managed to drag it out. [clears throat] So, let's go over some of the Google IO. Well, I don't know whether it was Google IO announcements, but there was a lot of Google announcements. So, let's just go through them one by one. Take a deep breath and pick them apart. It's too much of an intro. Uh, Google Gemini hits uh with 900 million users and goes fully agentic with Gemini Spark. We will get into that. A personal agent that manages your digital life alongside major UI overhaul called Neuro Expressive and new omni model for text to video. They released a lot. So let's dive on in. Um the Gemini app becomes more attract. Let's just go through these actually. Look Gemini 3.5 flash. We'll cover that in a second. Right. New model, fast model. Um welcome it. Neuralex expressive, a vibrant, dynamic, and completely reimagined design language for Gemini. Gemini Omni, our new model that can seamlessly transform text, images, and video prompts into cinematic, high quality video outputs. Daily Brief, a new agent that gives you personalized morning brief and organizes exactly what you need to know to start your day. Gemini Spark, a 247 personal agent designed to proactively manage tasks and help you navigate a digital life. and a Mac OSX app which I've just downloaded actually um which will come with Spark a little bit later on. So what is neural expressive? There it is. That's neural expressive. I preferred the old design to be honest. I mean, it's only it's only a little bit different, but I really don't [laughter] think they're really uh, you know, scratching at I don't know, clawing at I don't know what the phrase is. Clutching at straws a little bit for news here. The interface now features fluid dynamics, vibrant colors, and new typography with haptic feedback. I do not think this is vibrant colors. It's actually very dark, right? I mean, this is dark mode, the light mode version, which um which is why I was looking at it's it's now live. If you've got Gemini app, you can use it right now. Um but I wasn't all too excited about it. Let's switch [snorts] on uh personal mode. So, you got the color. As I say, I have my phone in black and white. You see that? Whatever. Let's move on. Uh, Gemini Omni. This is an interesting one. Turning your ideas into cinematic videos to unlock your creative potential. We're introducing Gemini Omni. Let's actually click on that. I don't think I clicked on it. A model designed to turn your imagination into reality by seemingly combining text, images, and video out inputs. Gemini Omni allows you to generate stunning high-quality video outputs effortlessly. Now, the interesting thing is how is this a model? Oh, I hope this uh let's see if we can get rid of this music. Can't even Oh, here we go. There it is. Okay. So taking video of the new person but also videos generate because they had the other I mean they've had VO which is another like all of their video products so use um do workout details. Aries, a great way to edit videos before app, you know. Uh, lastly, now a banana. Edit your videos through conversation. I want to know what the what the UI looks like here. Prompt. The lights of the apartment start turning on in sync with the music. So, they've obviously provided [music] the music here. Turning the light on. The dim lights in the room put a black and white checkerboard room inside a glass sphere that floats tracking above the hand. Oh, that's a complex uh I mean it's cool though. Some weird artifacts in the background there. But um what I did see though was Marquez Brownley showing the unboxing video. Gemini Omni made this, right? So it's an un It looks really really good. Like the technicality looks very very good. But as Marquez has said here, it defeats the whole purpose of watching an unbox video. Now, I get what you're saying, but at the same time, it's like you you you start with the with the what it can do and then it's up to creatives to actually bring an element of creativity into it, right? Maybe it's a small element of a bigger larger edit like this is you wouldn't release not saying that you're going to release this and this is like one and done. I think god you know I think that's the lazy route to think about AI. I think you haven't I think you haven't played with AI long enough if you think one prompting is the, you know, what you get out of one prompt is the thing that we're going to use AI for. I mean, there will be people who use AI for that to be fair, but you know, [snorts] um, so yeah, pretty pretty cool. They again, I wonder how it competes with uh, VO and I think there was another one. Forgot what it was called, but like how all this plays together, it's getting very confusing, Google. Very, very confusing. Gemini Omni begins rolling out today to Google AI Plus Pro and Ultra subscribers. Well, right. So, I actually I think I actually have this. I should have uh prepared a video. There it is. Maybe we can do something. Let's Let's add a Do I have Maybe I can go I mean, it has create video. Add an image to your prompt. Let's Let's record a video and add it. Let's see what we can do. Tune in to Command AI where we talk about all things web design and dev in the world of AI. Okay. Should I Should I put me in space? Should I put me in space? Traveling through a wormhole. Okay. Put me in space traveling through a wormhole. Lots of echo and sci-fi sounds in the background. Okay. I can imagine this is going to take a little while. So, let's creative with create with Omni. There we go. So, let's do animate. So, there's my UI there. And I'm going to select anime as the style and hit that off. And off it goes. Let's pick that up in a sec. It's thinking as well. I can see refining the prompt. I'm generating your video. This could take a few minutes. So, check back. See when your video is ready. Nice. Daily briefing. Bit of a yawn fest this, but they add daily briefing in chat GBT. Do you use chat GBT? Daily briefing. Is this a useful tool to you? I don't know. We're introduc in introducing Daily Brief, an agent that gives you personalized morning digest that is designed to be your first stop every day. Built on the success of our recent Google Labs experiment CC, Daily Brief gives you seamless, intuitive entry point in the world of AI agents. Now, I can see this why this might be helpful, but you need to use Gemini to get the best out of it, right? Or does this plug into your Google history? Does this, you know, this is all Google at the end of the day? Is this a whole ecosystem thing where actually you start to benefit from it right away? Now, if I I don't have I I don't have uh daily briefing in there right now, which is interesting because Okay, video is ready. I mean, it's pretty cool. It's finished. I'm just downloading it. Engaging maximum thrust. Commencing wormhole jump. Engaging maximum thrust. Commencing wormhole jump. So, it didn't take my audio at all. It looks like me, which is Well, I say it looks like me. It's a It's an anime version of me. It's me and my t-shirt and my hat that I'm wearing. Um, but it doesn't didn't say what I wanted to say. So, let me let me follow up with that. But you didn't but you didn't say the thing that I was saying in the video. I mean, it's switch back to flash right now, but interpreting the critique. Sorry, you're just generating video. Check back. Let's see what it does. Uh anyway, back to daily brief. Yeah, let's have a little look. Once you opt in, Gemini works across your connected apps in the background. What are the connected apps? What are the connected apps? It gathers urgent updates from your Gmail inbox, tracks up andcoming events from your calendar, and compiles regular relevant follow-up details and skimmable briefing. [snorts] I don't know. I don't know about that. Goes far beyond a simple summary. Daily brief actively organizes your priorities based prioriti prioritizes prioritizes prioritizes based on your specific goals and even suggest immediate next step. So it does help to use Gemini to do this. So yeah, I don't know. Okay, here we go. Daily brief rolling up today to starting in the US. I don't get it unfortunately. I would like to try it, but I'm in the UK as we've established. This is the interesting one. So, Gemini are going to be taking on Open Claw, which I don't know, given their Google, they could be on to a winner. Gemini Spark from information to action. We're also introducing Gemini Spark, a 247 personal AI agent that helps you navigate your digital life. Spark represents a big shift for Gemini, transforming it from an assistant that can answer your question into an active partner that does real work on your behalf, under your direction. Now, we're going to talk about the Gemini CLI because if you think about this, this needs a computer. This needs your information. This needs your apps. Gemini Spark runs on Gemini Gemini 3.5 and uses the anti-gravity harness. So that'll be something local. It's deeply integrated with the workspace tools you rely on daily. Gmail, docs, and slides are more even better because it's cloud-based agent. Yeah, see Spark keeps working in the background even when you close your laptop, lock your phone. That combination means Spark is ready to take on complex tasks off your plate so you can be more present for what matters most. or as we said about earlier, people are going to be walking around with their bloody laptops, not even laptops open, but their phones bloody sending off requests to their their agents. The interesting thing is, we'll get into that. Actually, with Gemini Spark, you can set reoccurring tasks with triggers, automatically piles monthly credit card statements to flag new or hidden subscription fees. As an aside, um you can't do stuff like that with um Claude. I I tried opening it in my bank app to kind of start helping me do stuff. It it refuses, which quite interesting. Obviously, I can work around it and maybe give it files or whatever, but yeah. Okay, the video is done now. Commencing wormhole. But you didn't say the thing that I was saying in the video. Did you hear that? But you didn't say the thing that I was saying in the video. Okay. Well, it looks cool. Let's try and give you that. It looks cool. I'll give it that given it kind of looks like me, but really not because it actually has hair. Uh, teaching new skills director to check your inbox for ongoing updates and create complete workflows, blah blah blah. I think we get the picture, right? [snorts] The thing about this is is the um I made a video on hostingers implementation. I also made a video on kilo codes implementation on a cloud-based openclaw agent. Now I think they're good for a they are good for a lot of people but because they're limited to the infrastructure that it's running on you can't install things like let's say image magic right I do a lot of conversions I don't do it with open claw mind but um or whisper which is a which is a transcription tool uh which transcribes videos and audio and stuff like that they're all applications that are installed on my Mac that I can just run through open claw like with home brerew and stuff like that but you can't do that in these online versions you know um obviously you can set it up to do that you can you know whatever um it's a lot more involved like you have to have your own server SSH in and do this and that Google aren't going to let you do maybe they will but Google aren't going to let you do that uh kilo code don't let you do that uh I don't think hosting will let you do that, but uh well, they do, but again, the one click setup, they don't let you do that. It's just going to be limited to a lot of Wow, here you go. It looks like it's going to be limited to a lot of very, very useful tools. I won't they won't lie, but limited nonetheless. Gemini Spark will roll out to trusted users this week, and we're planning to roll up beta users for US Google AI Ultra subscribers next week. So, if you're in the US, give it a go. Let us know. Gemini Mac for Mac OS. Take control of your desktop. We're working on big updates on the Gemini for MacOSS. Right. This these aren't in there just yet. [snorts] Gemini Spark to the Gemini Desport. So, this is what we were just talking about actually. So, they're going to bring it to the desktop app, but they're going to start with the the the online version um first. So it can help with task involving your local files and automate workflows across your desktop. We're also innovating on a new voice experience on new voice experiences in Mac OS app similar to what we previewed at the Android show. What did they preview at the Android show? Um, voice to text. Uh, turn spoken thoughts into polished text. Gboard. Okay. So, it's just a Yeah. Removes and Rs and things like that. Um, there we go. Us and Rs. Or what about uh that happen as you think aloud? Uh, using the context. Uh, I always say um and things like that when it's like when I can't think of the thing to say, I say things like that. Sometimes it doesn't make any sense. And sometimes I've said something that, you know, you can't possibly imagine what what other things are and yet I say things like that. I need something to tell me to or at least edit those out. That can happen as you think aloud. Uh, Mac OS uh MacOSS app is available to download right now for all users. And I did. I downloaded it. I opened it up. It's literally just right now. It's just the same as the mobile app. It's nothing nothing exciting just yet. Uh but yeah, making strides. I think uh I welcome it. But that's kind of all they released on that side of things. Let's dig into the other model side of things. Uh let me just double check. Gemini, we covered that. covered this. So, anti-gravity 2.0. Now, the first one was abysmal. Let's see what they've got in store for the next version of anti-gravity. Anti-gravity 2 is a standalone desktop application that fully delivers on truly agent optimized experiences. And we need to we need to talk about this like it was the first to really bring about this agent view. might not have been the first to be honest, but they they made this they made they released it when it was like pretty much brand new. There might have been some smaller um studios building an app that has like an agent view but they did it first seemingly. [clears throat] Um users interact with powerful agents both synchronously and asynchronously and there is no IDE dun dun while remot retains many of the core principles of the anti-gravity's ID agent manager in the surface. It is completely separate desktop application available to enterprise powered by the latest a oh is this not is this not something that I mean they're all looking the same aren't they really? They are all looking the same. Um, we'll get into the open source stuff, but yeah, I wonder if they all open source this. Let's walk through a high level the new features as demos, deep dives. Check this out. Okay, let's do that. Sub agents, asynchronous task, JSON hooks. That's cool. I mean, you know, Claude has hooks, but it's quite nice. Agent management revamped to allow for easier sorting and identification. schedule tasks, which is nice. Uh voice slashcomands grill me. Uh goal tells the agent to run until the specific task is completely finished. I think Claude recently introduced that as well, which is quite nice. Some only good at the prompts you give them, and it's common for users to miss important details or to overlook providing critical guidance. With this command, the agent will ask for clarification questions back to the user line on specific. It's quite nice browser. Most beloved features of that was Yeah, it was another thing as well. It was one of the first to introduce like a browser connection through like a Chrome plug-in. Uh autonomously launch and uh actuate and observe browsers. Kind this bloody thing. So, here we go. Um these slash commands here. Goal grill me. And that was kind of it to be honest. Uh [snorts] let's go back to here. These agents are more powerful than before. [snorts] And and that's the thing, you know, I saw flash 3.5 flash there. Can you see that? I can't because they've hijacked my scroll. I can't zoom in. 3 Gemini 3.5 flash there. You know, I mean, apparently it's only a little bit worse than Well, no, that was composer 2.5, so ignore me. But, uh, it was it was it wasn't doing too bad if I remember rightly for the, uh, Sweet Bench. Now, I'm going to read this cuz I think they're bringing out the S the CLI as well, the anti-gravity or the Gemini CLI, which is a really IDE looking for. Here we go. These CLI Um I I do like working in CLI to be honest. Even though I do sometimes work in the cro in the clawed desktop app and work in claude code in there, most of the time I just prefer the CLI. I just I don't know. I think it's a focus thing, clarity thing. I'm not too sure. But what I think this um has created a bit of fear on is that they the um I think it was the Gemini CLI was open source and a lot of the other you know Chinese um uh labs built their CLI on top of Gemini CLI but if I remember rightly reading somewhere they've discontinued it. At Google, we will drive the most long-term value to our users by having a single cohesive developer product offering the leverages and shared harness co-optimized with Gemini models. As part of this unification effort, anti-govern CLI took inspiration from core Gemini CLI product and harness components. This brings the best insights and focused CLI and agent experience for developers. Read more about migrating your workflows from the Gemini CLI to the anti-gravity CLI. The anti-gravity CLI will continue to inherent incorporate improvements in the core harness and agenda paradigms. Anti-gravity 2.0 will be optimized for comprehensiveness in its feature set while the anti-gravity CLI will be optimized for speed and low overhead. Interesting. So, they've kind of established or or admitted that there is an overhead to running this, you know, UI, which makes sense. [snorts] Um, yes. I read somewhere they're going to be discontinuing the uh Gemini CLI or not discontinuing it, but maybe um not making it open source. I'm on this blog post by the way. I went over Yeah. Like why would they Okay. Gemini CLI will remain access by pay Gemini and enterprise. So they are push it looks like they are pushing you towards the Gemini CLI and they're probably going to discontinue the sorry the anti-gravity CLI and they're going to discontinue the Gemini CLI which is a shame. Which is a shame but they've got an SDK. If I share this one they got an SDK as well. But all this being said, I did a real massive consolidation of all of my um Google Cloud payments that going out because I I was getting like three random transactions and like four random transactions of only like a dollar or like a few cents actually a few dollars some of them. It's like what is all happening? Like where am I paying for Gemini? Because I literally have no idea where I'm paying for Gemini. It's not a lot cuz I don't really use it all that much. I use it a lot on my NA workflows just because it's easier to set up and cheaper. But yeah, who pays for Gemini? Are you paying for Gemini? I don't know. But yeah, that's it. That's most of the updates to the anti-gravity um let's call it ecosystem. Um obviously Gemini 3.5. Do we want to go with Gemini 3.5? Do we want to go over that? Um, sorry, I'm looking through the blog posts here. That's cool. We've gone over that one. Let's Let's go over Gemini 3.5. So, amongst the slew of other products that Google released at Google IO, they released a new Frontier model, the Gemini 3.5 Flash intelligence at four times the speed of other models. Let's have a look. Um, today we're introducing Flash. 3.5 Flash is available today to billions of people globally. I saw it on my little app there. So, that's cool. Uh, we're also hard at work at 3.5 Pro. It's already being used internally and we look forward to rolling out next month. So, looking at the benchmarks because of course we've got to look at the benchmarks. [snorts] It's doing all right. And especially if you take note of how well it's doing against 3.1 Pro. 3.1 Pro, I can already see. Sorry, I'm giving my leg a scratch here, but can already see terminal bench 76.2 whereas Gemini 3.1 70.3 bench just over Gemini 3.1 Pro. [snorts] I did not have a good time using Gemini 3.1 Pro using on code. So given it's a little bit better and Opus is the the leader there, you know, I think it'd be good for simple tasks, but yeah, MCP Atlas 83.6 is winning there. Tool, it's winning there. Never heard of those though, to be fair. [snorts] OSWorld GT5.5 is winning there. Finance agents quite interesting. I thought Opus would be the winner there. Uh where are we? Opus 51.5 57.9 GDP vow 165617453. This is the one that I am in I'm following. Didn't know GBT 5.5 was so good at that. But we won't bore you with the benchmarks, but it's good to have a little nice idea about where where it's going. Um, here is a little diagram of the intelligence versus speed. So, speed is nearly 300 tokens per second if on the right hardware obviously. Now, it's not going to beat 3.1 flashlight. Uh, but that's a tiny tiny model. This is the intelligence index here. So, it's dumb as [ __ ] Dumber than HighQ, but on par apparently with intelligence wise, on par with Opus and and bit better than 3.1 Pro. I don't know. Uh Aentic scales at aentic tasks at scale. The balance and speed of performance 3.5 uh flash ideal for tackling long horizon agentic tasks. what used to take a developer days or or say weeks, 3.5 flash can now help complete in a fraction of the time. I would assume that you would need a big old model um to set out the plan and plan it out, but then a uh 3.5 flash to kind of do the, you know, the the lifting the the dog's body work. Basically, when coupled with updated anti-gravity harness, 3.5 Flash becomes a powerful engine for deploying collaborative sub aents to tackle problems at scale. So, there you go. sub agents being handed the context, being told what to do. Um, it will just churn it out. Real world impact personal AI agents 3.5 flash. 3.5 Flash is available today. And that's it. That's it. What a what a what a whirlwind of products from from Gemini. I think they uh I They released a few physical products as well. Um some glasses, some you know few things here and there. I uh with some varying degrees of people's uh approval on those on those uh who's that Schmidt guy? Is he part of Google? Did he make a speech recently where he just literally got booed because he was undermining all of these students at their graduation telling them that AI is going to take over and stuff? It's like, hang on, you're in front of a bunch of people about to enter the big wide scary world that AI is scaring everyone and you're saying that it's like they need to be empowered to say that you will take charge of AI. You will be I know what he was trying to say. But it didn't come out like that at all. He was trying to say it's an AI world. Like take advantage of it. It's a it's whatever. It's very empowering. But it came across like like the AI will be the thing controlling them. They are at the helm of AI when it should be the other way around or at least giving that impression, you know. Anyway, um yeah, Schmidt, what is it? I don't know who he is. I don't know. Either way, it wasn't a great It was a great week for Google, not such a great week for the responses from Google, but it is what it is, right? Let's go through. We've got three more. Um, Grock 4.3 powers Open Claw, the update with sturdier plugins and messages, messaging fixes. Open open claw 26.5.2 drops with XAI's Grock 3 4.3 support enhanced plug-in stability leaning gateway and agent paths plus major fixes for Discord, Slack, Telegram, and WhatsApp integrations promising less downtime and more reliable AIdriven workflows. Let's see how Grock managed to worm its way into open claw. I think I I did a bit of research before the show and they were I couldn't find what it was, but I feel like Grock made a massive like, you know, entrance onto the scene with some integrations with a bunch of stuff. Um, I think I think it was open code. I think it was open code. Yeah, because there was that. Let me find it. There was a video made with Omni and um uh featuring Dax from Open Code. I'm just on Twitter right now. Uh dot dot dot maybe maybe he did retweet. I don't know. He tweets a lot. Uh [snorts] uh. Now I can't find it. You know the um you know the sexy avatar that Grock used for their sexy chat app? Well, it starts off down there. I think it either I started off down there. My eyeballs started off down there. Her legs, you know, pinups and all the rest of it. But it's Dax's head. This dude with a beard. Yeah. I can't find it unfortunately. Uh, but it was funny. Anyway, I'm getting distracted. So, Grock, use Grock in OpenClaw. Use your Super Grock X premium inside Open Claw and open source. So, I've never rated Grock for anything. Have you been using Grock for anything? Is it good? Is it fast? Is it intelligent? Is it underrated? Is it overrated? Um, start today. Log in and use your super grock uh or x premium subscription inside openclaw. Open claw is an open source local first agent. I don't think you need to know you know we all know what openclaw is a raspberry guy persistent memory across systems. Openclaw connects to WhatsApp. If you already ever got premium yes we get the picture. Okay. So you install it you on board yourself. You set the or choice as XAI and that's it. Okay. Well, that's kind of it. Um, I think the the thing is here is that just Grock are I think well my thoughts are that I think Grock are trying to get out there and make a make um a name for itself for us to actually start becoming a Here we go. Hermes as well. Skills. Grock skills. Grock build better. Look at all these like Grock connectors. This is all within the last few weeks. I don't know. I feel like they might be doing something and maybe I need to pay maybe we need to be paying more attention to Grock. Uh but it is Elon Musk. So [snorts] here we go. Open code. Anyway, yeah, that's all I've got to say on that one. So yeah, let us know if you're using Grock and uh get jiggy with it. So we have clearing up all of my This will This will be an article that I've not looked into just yet. Perplexity deployed a new extrative compression model that surgically removed distractors from retrieved context. Now this sounds suspiciously like composer 2.5 training. Keeping only the evidence needed to answer a query cutting the noise latency and cost simultaneously. Also sounds like the subq stuff we went through next week. Let's go into it. Let's get into it. Make this bigger. Queryaw aware context compression for better snippets. Improving the quality efficiency frontier model context through query. Okay, repeat yourself. Every answer starts with evidence. In a search system, this evidence usually comes from documents, titles, summaries, chunks, snippets that are passed as context to an answer model or an agent. The raw [snorts] context is noisy. For instance, a web page may contain information needed for the user's request, but it also may contain distractors like thought it would be the same thing. Uh these distractors can appear both in the main body and in ancillary elements such as nav text elements. Irrelevant ads. So it's specifically talking about websites here and we spoke about the unreasonable effectiveness of HTML last week might be relevant. First, it hurts accuracy. Even the most powerful models have finite capacity. Long streams of irrelevant information waste model capacity, resulting in context rot that impairs a model's ability to address the actual user request. We get it. Um, blah blah blah. Introducing Okay, we get it. It It's okay. So the same they heavily impro they heavily invested in context position. One way we do so is refining our snippet generation algorithms. We treat a snippet generation as context compress context compression problem. Models looking for specific information do not benefit from generalized summaries or experts of a document. Rather they need the smallest most surgically extracted piece of source information to respond to the user's request. Everything else must can and must be removed. Okay. A good snippet should therefore do three things. Promote accuracy by presenting the precise evidence needed to ground the model's response. Reduce latency and cost by removing irrelevant context. And ensure traceability by preserving citation fidelity. This month we deployed a new state-of-the-art snippet generation generator within our application's API platform. Let's open that up in the background and we'll have a little look in a minute. Core of this module is query array context compression model which decides for each query and candidate result the subset of spans. So it seems like it seems like it's an extra request that summarizes then summarizes again but maybe I'm being too simple with that. First order business define the task itself what we want the model to do and how we should define input and output schema snippet methods. So, actually, let's just dive into the API. Uh, it's just a link to the actual API. Search embeddings sonar. Maybe it's in the search API or was it called contextaware? Uh, query optimization. No, no, I'm not too sure. Let's let's dump back into here and see what if we can learn something more. Snippets methods fall into two broad categories. The first selection based find a relevant window chunk sentence or passage pass downstream. The second is generation based ask model to write a query focus summary from the page. Okay. Well, rather than getting too technical and let's have a look at this. because it's all well and good showing it. Okay, we're going to keep that. We're going to drop that. But like, how does it do it? Do you know what I mean? Again, that's probably just way above my understanding, but it's cool. Anyway, it's always nice to see um improvements made. I'm kind of a fan of Perplexity. I do like I'm using Pexity browser right now this second. Um it's interesting what they're actually they're actually focusing on web and search which is their wheelhouse. So it's good to see. Yeah. I mean let's celebrate wins, right? [snorts] Okay. We're going to wrap up on this one now, boys and girls. Figma launched its own AI native agent that runs directly inside a collaborative canvas. Natural language prompts to generate designs, edit existing ones, or automate iterations with multi- aent support and claw codeex integrations already built in. So let's take a look at their little video designer and also work on agents. And today we are so excited to introduce to you to new agent. Let's jump into a few examples of how you can put the Figma agent to work. First up, there's some task that just imagine component that needs to be swapped. So, you'll see she's describing a change uh set all the the component. So, she's mentioning a component there. What did she actually say to their higher states? So, that's probably something a very domain specific uh request from the Figma agent there. Right. Find all text larger. Let's pause it on that. and change the font to enter. It's a very very contextaware, very not um like just helpful. They're not they're not leaning into the design aspect, which is interesting. Here we go. Add code syntax for all web for web to all my variables. the name should be. So, so you're actually able to control the file itself fairly under sort of not the code creating the file, but the the the structure of the actual Figma file itself, which is really easily uh which is really nice. And now they're getting into like design, give me three design options for this organic go ahead and uh design those files. And it looks like it's related to the original one, which is quite nice. Um, but that's where you're not going to get the best out of AI for design to be honest. It's those generation things, but useful nonetheless. And here, styles and nails. This looks like it's that says design library. What does this generate? Okay, so that's gone ahead and generated a bunch and it looks like it's taken advantage of all of the Figma stuff. So, it's definitely tuned to Figma. Oh, yes. And also, they're going to go through comments here as well. So, you're going to sort these comments by theme. Creating a new rev of my revision of my design. So you can actually use all the comments to create a new design instantaneously. That's really nice. That's really really nice. Really nice actually because we've just handed in a design and they've they've left loads of comments on the Figma file for all of the copy changes. literally just use the AI agent to like uh add all of the comments, all the text changes to these and highlight any that don't make sense or whatever. Amazing. That's really good. And that's kind of it. And the interesting thing is because I got a text about this as soon as I woke up yesterday morning. So, it's brand new off the press. But right before, this is crazy. Right before I received that text, I I tweeted this. AI is not there for designers. But what you can do is use it to research, iterate, explore, draft, implement your idea. We need a pro designign tool that supports these areas. Then you might just win over designers. Now, Figma didn't necessarily uh doesn't necessarily support especially the research or um well, yeah, the research um aspect of this actually. They do everything else. So, I think they've really approached this in the right way because designers faking hate AI. Designers faking hate the fact that uh their jobs are not at threat. They really aren't like you you the an AI that does design will satisfy people who have no idea about design. And that's okay. There's nothing wrong with that. Like I've AI designed a bunch of my apps and I'm mostly okay with them. But a better designer could probably look at them and be like, "Eh, you know, they might follow basic UX trends, which is where my focus mostly is." Um, but AI design just isn't there. And I don't think it's ever going to be because, uh, I spoke to Kabaza about this yesterday. Like there are two types of designers. Some of them are very, you know, methodical and spreadsheet designers, let's call them, who, you know, design based on, they're probably more UX focused designers who who don't go too crazy. But then you've got the designers who literally want to put their soul into the design, even if it has been done a million times before or it's whatever. They want to know that the thing that they're is the exact thing that's in their head goes onto a piece of paper, metaphorical piece of paper, and the chances that AI is going to do that is like slim to none. So those are more artistic designers. And then obviously there's a scale in between, you know, but uh you're never going to satisfy that half, especially the latter half of designers because they don't want AI to to do it for them. But they shouldn't they shouldn't rest on design, you know, they shouldn't um sorry, they shouldn't rest on AI. They need to be uh embracing research using design not even doing it in a design tool just staying in chatbt or using um perplexity or something like that researching iterating. So once they've designed something using it to and and and not taking the whole design and iterating on it. Maybe it's just a button, [clears throat] maybe it's just a section or something or it's a font choice. Using AI to just throw out some of that and and leaning into this idea of exploration, you know, whether it's give me some random ideas that I would never think of, just just spice this up a little bit. What do you think? and then you might reject them or implement this font and this font and this font in three separate designs of this or three separate iterations of this exact design. So it's only changing the font. All of this stuff looks like exactly what Figma design agent is doing drafting as well. you know, again, maybe your prompt is megalong. Um, but you can just get it to spin up, you know, especially if a design has been done a million times before. Like how many times, what if you get a client that just wants the very basic um dashboard app, you know, that lefthand side uh panel um you know, the sidebar um you know, widgets, it's been done a million times. Get an AI to spin that up. Like you don't not every single design uh needs to boil an ocean notion. Not everything not every single design needs to reinvent the wheel. Like let's let's leverage patterns when appropriate. But if you have got artistic freedom then that's when you can bring in your own creative touch right. and then implementing your ideas similar to the drafting which is like this is exactly what I want just go and do it for me like I know what I want and and I AI might be able to get 90 98% of the way there and then it's in your Figma file then and then you can just change it up make the make make the border radius nine instead of 10 or whatever do you know what I mean like just use it as a starting point while you set that up whilst you go to the toilet you come back most of the work's done that would have taken you [snorts] like 45 minutes or whatever to get to that point and then you just come and just tap it in the net, you know, that's a football reference for scoring a goal, you know. Um, so yeah, I think Design Agent covers these uh and and the they were very careful in their video to not [ __ ] on designers, to not take away designers because their their role is not, you know, someone who who lives in Figma, their their job is not going anywhere for a long long time. you're always going to get people who still want a unique and crafted um piece of work that AI will never achieve or at least you'll be prompting for ages until you know you get it. The good thing about AI design though is that with well I guess the worst thing about AI code is that you'll have hallucinations somewhere. Um, there'll be code that's written that just doesn't do anything or it will leave code in that's like irrelevant anymore. Good thing about AI design is you literally you're looking at what it created and if it's gone something wrong, you just correct it. I mean, it's a it's a lot more it's a lot less complex than uh code from a technical perspective. But either way, um, yeah, at least you won't at least you know when a hallucination has happened, you know. So yeah, Figma design agents, I'm a big fan. I think this is the right step in the right direction for for designers to begin to embrace AI instead of rejecting it because it's not trying to replace them. It's not trying to stick it in the way of of something they they they want to do, they love doing. So I dig it. Right, boys and girls? all 37 of you. I'm very glad you guys tuned in with me on my own. Sorry Kabaza couldn't be here, but you know, life has to happen outside of the the internet, you know. So, [laughter] um that is all we've got time for. We It will be me solo again next week unless I decide to take a break because it's very exhausting doing these streams on my own. So, um but we you can join us next week. We've by the way, we've just launched a website commandow.com. Go check it out. Go let us know there any typos because I think Kabaza handed me the design but I think that was mostly AI generated based off of a skill of our branding which was not AI generated but that website was AI generated and I spent all yesterday or day before getting it. So that website is up and running. Go check it out. Let us know. Um, I don't think you can contact off off it, but you can email me at samuel@commando uh commando.com and let us know what you think. If there's any bugs or anything like that, would really appreciate it. Or if there's something you want to see. Also, check out insiders.commandio.com. We're collecting um, in fact, I'm hoping it's still up. I'm pretty sure it is. It is. Let me share my screen. Go to insiders.comishai.com and fill out this form. We're really interested in uh creating a community and it will be an insiders community. It won't be for everyone but for those who join we want to really help uh empower them bring them up with their AI knowledge and we want to know if this community was to take place what would you sorry I should have uh showed you what would you want to learn about as a little thing here is it AI tools and productivity is it building apps design we don't know I don't see my topic you can fill out your own thing here right again this is AI generated from our skill really cool. Um, yeah, let us know. I should probably put a popup or something on the website that we'll link to this. But yeah, that is all we've got time for. Check uh tune in next week. Loved having you around. Um, thanks for tuning in. Wait for the clips for next week. Um, and until next time, keep on vibing.