|
|
|
Zero full-attention layers, MIT license, native 1M-token context
|
|
Beijing startup NaiveAI open-sourced the MIT-licensed Naive-N0.5-Flash on September 27: 309B parameters, zero full-attention layers, a native 1M-token context, per NaiveAI's model card. Price it against your current long-context coding-agent bill this week; self-hosting needs roughly 315GB of FP8-capable GPU memory.
|
|
|
|
|
Stanford ran a humanoid robot with no trained control policy.
Stanford and Caltech's HomeBody lets GPT-6 Astra call a five-skill library directly, and a Unitree G1 tidied an unfamiliar kitchen with its perception stack on a single RTX 4090 laptop, per the project page. If the approach generalizes, robotics teams maintain one skill API instead of training a policy per robot, and upgrading the reasoning model upgrades the robot. Stanford lists the current costs itself: pauses between skills from Astra's latency, and finger servos that overheat on long runs. See the five skills Astra calls.
|
|
Tool to consider · partner link
Self-hosting a 300B model needs real infrastructure.
Naive-N0.5-Flash needs roughly 315GB of FP8 GPU memory just to hold the weights, per NaiveAI's model card, so most teams will call its hosted API instead of self-hosting it. If you're standing up the app or internal dashboard that calls that API, Cloudways handles the managed server, patching, and backups so you are not babysitting a VPS. Spin up a Cloudways server.
|
|
|
|
|
Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.
|
|
|
Affiliate disclosure
·
Pondero earns commissions on some links. This does not affect our editorial picks.
|
|