—
TIMESTAMPS:
0:00 Kimi K3 is destroying the benchmarks
00:51 It's the biggest open source model yet
01:13 Releasing weights soon
01:26 Accessible by Kimi servers only right now
01:43 Prices for Kimi coding plan
02:06 Beats GLM on its lowest setting
03:02 3D generation skills
03:44 Still not good enough beyond hobby projects
04:45 API pricing comparable to Sonnet
05:36 Multimodal capabilities are a gamechanger
06:40 Almost as intelligent as Fable
07:09 Cost per task
—
Unlock the full potential of your online presence with Kabarza and Samuel—experts in web design and development (respectively), powered by cutting-edge AI solutions. We blend creative design with advanced tech to deliver smart, high-impact websites that stand out. Ready to elevate your business? Contact us today and see what AI-driven innovation can do for you!
LINKS & RESOURCES:
Website: https://cmdaishow.com
Check out Kabarza's amazing work: https://kabarza.com
Visit Samuel's website for more: https://samuelgregory.co.uk
📷 Follow on Instagram: https://www.instagram.com/cmdaishow
—
HASHTAGS:
#kimik3 #moonshotai #fable
#ai #podcast #aidesign #aidevelopment #vibecoding #webdesign #webdevelopment #ainews #webnews #designnews #devnews
Transcript
Probably less than two weeks later, Kimmy, the other kind of Chinese open- source competitor, lands Kimmy K3, which really does blow Fable out of the water. Command AI. I mean, to be butting up against Fable and GPT 5.6 on on Deep Suite. This is, by the way, a closed source benchmark. So you can't benchmax against deep suite because you don't know you don't know what the test is, right? Terminal bench is beating Opus 4.8 and just just nipping at the heels of uh 5.6 Soul. It's winning on program bench, which is crazy. Uh same again here. It's nipping at the heels of of Fable 5. These are crazy benchmarks. It's a 2.8 8 trillion parameter model, which is probably one of the biggest open- source models I've ever seen. So, when we talk about open source and running locally, you need a lot of lot of hardware to to run this one locally. You're definitely going to be waiting for the quantizations of this model. Uh, which they haven't released yet. They haven't released the weights yet. I think it's coming January 27th or something like that. uh July 27th. So in 10 days they're hopefully going to be releasing the full thing. Um you have access to it right now through Kimmy uh and Kimmy code which is quite a nice CLI actually I have to say which means that you're always going to be going through their servers. So again something to bear in mind. These are Chinese servers. They can see everything you're doing and these are the sorts of prices we're looking at for those um coding plans. However, once they've released the weights, this is when other people can host it and probably be able to host it on US or European soil. Um, and well, let's see what the prices look like when they're on there. But it's definitely an amazing feat of engineering. Um, looking at these these uh benchmarks here, like there's GLM already the lowest setting there already beating GLM at roughly the same price. And yeah, I think this is an insane insane model. I think this is a And there's Inkling there. Look, just coming out. They were head to they were neck and neck there and then boom came straight out of the gate and just completely killed it. But it is 2.8 trillion parameters. So this is a third of the size. So you know the end of the day the more you put into something the more you're going to get out of it. So these are the open uh models here. used it. It's neck and neck to close to every like frontier model and in some tasks winning. It's insane. Here apparently it made these games and in this case it made even the models. That's insane. Mhm. Like they have an understanding of the 3D world and you see there are like many examples of like games. It's it's just getting out of hand what these models can do. And I I'm like, how can Open AI and Claude uh Anthropic make money? Like I mean, how can you justify it? Yeah. Like justify their entire existence. Yeah. I don't know. I've I've got one more day of Kimmy. I've been using it a little bit today. um getting it to solve some like light problems or something like that. I still don't have the confidence to give it like anything outside of hobby projects just yet, like the confidence, but I think I'll need to use it just a little bit more to see what I can actually throw at it and what kind of Yeah, what kind of work can we give something like this? Uh again, this is on Chinese servers, so maybe something to not throw your uh big open source project code base at at all. Again, just little shitty little projects are fine for now. And then we'll see what happens after July 27th. Will it be as bad as Brock? That's the thing, man. That's why that's so disappointing that that's they've already revealed themselves to be absolute scumbags by, you know, taking advantage of our trust. Kimmy K3 agents download update latest Kimmy app from your mobile store. You can download all that there. Um, I mean, I guess you can use it in the in the in the whatever. They've got their own chat interface as well. Um pricing is 30 cents uh in uh for sorry for a cash input. Uh $3 for a cash miss. So quite a bit massive difference between that and that. So once you've cashed it's pretty cheap. But the first one is $3. This is this is Sonic prices, I think, or old Sonic prices, I'm pretty sure, which is quite crazy. Um, but is it as efficient with tokens? Um, I'm not that this is also something to look into. Uh, I'm looking at their docs and it can also do like video editing particularly good. That's actually a really good point. That's a really good point. So, because it's multimodal, right? Yes. So GLM was great, but it wasn't m multimodal. And it is very uh helpful to be able to pass through visual cues on whether something's working or not. By the way, I I don't I'm not I'm not in the same camp as some people who are like, well, how does it how can it do front end if it doesn't have visual capabilities? Funny enough, you can interpret a lot through understanding code. You know, it's like when you look at something, but oh, that's about 20 pixels or whatever, like because you're looking at something. The reverse is true. If you can see 20, you can visualize things in your mind. So, I'm not 100% in the camp of like just because something doesn't have visual. I mean, like I say, it score uh GLM 5.2 scored better than Fable. Um, and it doesn't have visual capabilities. So, you know, proof is in the pudding there. This is the intelligence. This is kind of worrying in a way how close it is to fable and intelligence. Like this is this was the biggest thing about Fable is it intelligence I think and it's like there do you know what I mean? That's crazy. Speedwise not not so I want to I've downloaded uh GPTOSS. I do want to try this cuz it's always up there on the on the speed like unbeatable. But again third place here got electrocuted. Uh but you said something. You said cost per token, didn't you? Yeah, I just said I just asked if it's not cost per token, but per task if it's efficient with the token cost cost per intelligence index task. This is kind of it. So, it's like around 95 cents which um Claude took 275 and 25.6 took $14. So just a little bit cheaper than soul not cheap. So it's not Yeah. So not cheap. The tokens are cheap but it probably takes more token to achieve the same result. Yeah. However, yeah. More. Yeah. Yeah. Exactly. Cool. That's that's what I would say. So yeah, Kim Kim K3 I would um I do want to keep trying out, playing around with it. This is part of a larger conversation on my show, Command AI, which we stream live every single week. We discuss the news and all things related to AI in the world of design and web. Catch us next week and join in the banter. Command AI.