Agents
verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prioritize pretrain data quality enough that comm…
verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prioritize pretrain data quality enough that commoncrawl isn't good enough for you, you have to build a Whole
verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prioritize pretrain data quality enough that commoncrawl isn't good enough for you, you have to build a Whole Web scraper anyway, and if you wanna keep it current, you have to have indexing, and pretty soon you find yourself having built a total private low-frequency clone of Google as a SIDE PROJECT of pretraining, that you can then also reuse for the agent side inference. we do know that the labs do use third party search providers, but clearly this is one of those things where developing more and more of your own 1P equivalents is both a competitive advantage and an adversarial target for AEO Batesian Mimicry* that you will not want to share. * https://swyx.io/mimicry-reflexivity It's wild to me that both Anthropic and OpenAI have products that lean so hard on search, and yet they both obscure the underlying search index that they are using
Related
- microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale. this model uses zero synthetic data or dist…
- seems obvious now but 4 years ago i was getting blank stares when I talked about leveraging LLMs in VC
- Before enterprises can run with agentic AI, they need to learn to walk with their data
Source: Swyx (X) | 2026-07-31