Guillermo Rauch
The big lesson from AI is that everything is code. A slide deck is code. Design is code. That cool promo video? Code. Excel automation? Code. The universe? Probably made of code too.
Follow builders, not influencers.
The big lesson from AI is that everything is code. A slide deck is code. Design is code. That cool promo video? Code. Excel automation? Code. The universe? Probably made of code too.
Never a dull moment when you work at OpenAI. Absolutely incredible place.
Multi-model agentic systems clearly are the future. Great post by cursor. Their new research shows that a frontier model as a planner and orchestrator with a workhorse cheaper model can materially lower the cost of total tokens on a project, resulting in a 15X cost improvement.
“Few moments in a large task genuinely require frontier intelligence, such as the original decomposition, the design decisions, and certain trade-offs. Once a frontier planner has collapsed the ambiguity into a detailed, explicit instruction, less expensive models simply have to follow it.”
This is increasingly becoming the core design pattern of complex agents, because the tokens used on large amounts of work on many tasks don’t require the same intelligence threshold as the planning step. By routing to different models based on where you are in the task, you get much greater efficiency overall.
This provides the template for where the applied layer in AI will differentiate. You can only drive this efficiency if you know the domain well and have the ability to work with multiple model tiers; companies that can do this will across coding, finance, legal, healthcare, life sciences, and other critical spaces will be in a strong position to get larger workloads that would otherwise be too expensive for the customer to deploy.
The asset seizure tax will only make California more impoverished
SEIU-UHW should be ashamed of itself for the naked cash grab for their own personal benefit https://t.co/wfG9MyVpzc
Compute rules everything around me CREAM https://t.co/LkOQBYaDfu
Great reminder to stay frosty https://t.co/mJNqrKnYt5
If I were to hire today, this is how I would structure the interview process:
Round 1: in-person & no AI allowed; test domain expertise and knowledge on the fly
Round 2: a project that must be done with AI (impossible to complete without AI). Candidate assessed not just on the result but also the chat transcript with agents
Just reached 80k followers 🙏🏻
“all signal no noise no slop no hacks and still finds a way to grow her audience” is what I do here https://t.co/pI8h5ImgnC
There are basically two kinds of companies now. The ones built before coding agents showed up, scrambling to retrofit. And the ones founded after.
The second type is different from day one:
it's literally the greatest time ever to have product sense https://t.co/BkzV0sLwgP
The road to AGI is paved with economically valuable tasks.
That’s why enterprise AI is one of the most important frontiers. It’s where many of those tasks live.
Four years post peak web3 and crypto tokenomics debates…
turns out the tokenomics debate that matters is open vs closed weight, inference costs and model routing. https://t.co/THz3HEyx1d
this was a bug that was only live for a few minutes, ran into it on my personal account while doing some late night coding though
Many founders in the last 18 months are going to realize that just because there are no “moats” in AI left anymore doesn’t mean “scale and capital” become your main moat.
We have seen this movie before.. all of us will benefit from reading a little bit of history.
Webvan
Groupon
MySpace
Yahoo
AltaVista
Blockbuster
Nokia
(zombiecorns of 2021)
Each one was structurally and capital-wise in a very good position, but eventually got beat by a company with a much better unique insight. Or just collapsed under the weight of their own scale.
This is the needle founders have to thread: find a unique insight worth a 10+ year journey, while being prudent enough to not let capital & scale become a substitute for it.
This VC to “fast growing startup” train is insane to watch.. another three people just this past week.
I guess it’s similar to the IB/consultant to BizOps train I saw in 2012.
Top title for this cohort?
Special projects ⚡️ https://t.co/FqarRFC8Iw
https://t.co/HfFp5SlucN
we’re hiring a senior engineer to help work on our agent @every
must love agents. dm if you’re interested (for now preferencing for people I know on here!) https://t.co/RNbNiBQVoL
guys I’m basically Brian Johnson https://t.co/vgDGWhxBlZ
1) What https://t.co/kHWwdn0ZQs
Banning Chinese models will be the same self own as banning Chinese EVs.
Use one agent to do the work and another to review it against a rubric.
@trq212 explains how this works for video shorts:
“For something like, ‘Is this a good video short or not?’ you don’t have a deterministic answer.
You should have a separate verification agent read a rubric, review the short, and give feedback.
We call it self-preferential bias. When a model prefers its own output, it’s going to be more lenient at verifying it.”
📌 Watch the full episode here: https://t.co/mMrt4D5qU7
unironically this is happening right tf now https://t.co/8E0EI7Gq38 https://t.co/FxVx5B9jIi
very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction.
an open secret of "frontier" model training is that even without training on test, you can basically cheat by training on test lookalikes, enabling you to goalseek almost any benchmark number you want.
however when they are released open weights, 99% of the time the norm is that you do not get the datasets/rlenvs that would easily show you if someone was training on Temu Tbench, so there is plausible deniability. Alex and Omar discuss applying standard NLP distance metrics on hidden trajectories. There's no ultimate solution here, but they have some prelim explorations. It happens to support the finding that RLMs can generalize to unseen tasks that share latent structure observed in training.
https://t.co/74kfnYzuE9