We're building in Canada. 🇨🇦
We're building in Canada. 🇨🇦 For decades, Canada invested to build the research foundations that made modern AI possible. Now we have to build, train, and scale what comes next here at home. Canada's
Knowledge catalogue
We're building in Canada. 🇨🇦 For decades, Canada invested to build the research foundations that made modern AI possible. Now we have to build, train, and scale what comes next here at home. Canada's
We're kicking off our first cohort of Runway Builders with almost 100 startups building at the frontier of real time video with the Runway API – notifications going out to accepted companies today! If
We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensi
We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise t
we're taking an early bet on open models specifically, because they're SO much cheaper one point of reference: an app outputting 10M tokens/day costs roughly 250/day on Opus 4.8 versus ~24/day for Min
arXiv:2507.03373v2 Announce Type: replace Abstract: Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-ge
We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT. https
arXiv:2606.04233v1 Announce Type: new Abstract: A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. W
arXiv:2606.04989v1 Announce Type: cross Abstract: Although much is known about the physical danger of cycling situations, less is understood about the perceived danger of cycling. Furthermore, percept
What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmzbmr1 shared the story on the OpenAI Podcast with @AndrewMayn
arXiv:2606.04425v1 Announce Type: cross Abstract: Modern agentic systems transform LLMs from session-bounded assistants into stateful systems that persist and evolve shared world state across sessions
what makes building on replit different? it all happens in one place. → describe your idea in plain english, get working software → generate UI, add auth + databases, deploy → collaborate with your te
Enterprise AI is forcing companies to rethink where their most sensitive workloads should live, and private cloud is moving deeper into the center of that decision. For many organizations, the cloud c
arXiv:2606.04935v1 Announce Type: new Abstract: Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent
At Google Cloud, our goal is to let you run large-scale analytical and data science workloads with maximum efficiency so you can process big data pipelines, machine learning, and ETL tasks. We recentl
arXiv:2502.08870v2 Announce Type: replace Abstract: We provide an approach for the analysis of randomised exploration algorithms like Thompson sampling that does not rely on forced optimism or posteri
arXiv:2606.04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near t
arXiv:2606.04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target functio
arXiv:2606.04389v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise in psychological counseling, yet existing benchmarks rely heavily on highly cooperative simulated clients. We
arXiv:2603.09242v2 Announce Type: replace Abstract: The growing realism of generative models has blurred the boundary between real and synthetic content, posing significant challenges to reliable AI-g
arXiv:2606.04375v1 Announce Type: new Abstract: Differentially private stochastic gradient descent (DP-SGD) injects noise into every updated coordinate, making the injected noise energy scale with the
arXiv:2606.04361v1 Announce Type: cross Abstract: Age of Information (AoI) has become a central metric for the design of wireless update systems, especially in applications where fresh measurements su
arXiv:2606.04161v1 Announce Type: new Abstract: Different predictors often excel on different inputs, so picking the best one per instance promises higher accuracy than committing to a single model. I
arXiv:2606.04127v1 Announce Type: new Abstract: Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely v
arXiv:2606.04098v1 Announce Type: new Abstract: Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spli
When you burn so much money you run out of options… Anthropic co-founder and President Daniela Amodei said the high cost of developing AI models is driving firms like hers to look to the public market
Gemma 4 12B is the first medium-sized, encoder-free multimodal model capable of natively ingesting audio and video , recently released by Google. The model is available on Ollama with 11.7M downloads
arXiv:2606.05107v1 Announce Type: cross Abstract: We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tu
⚠️Why didn’t the hyperscalers wait until after the IPO in switching to pay-by-usage-charging? My guess is that *they literally could not afford to* —because it would bankrupt them. My reasoning: in “a
Why Japanese companies do so many different things: The internal logic of the world’s strangest corporations https://davidoks.blog/p/why-japanese-companies-do-so-many TL;DR > you have a firm that has
arXiv:2606.04662v1 Announce Type: cross Abstract: Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage rema
Agentic AI is making the semantic layer an essential enterprise priority because headless agents asking thousands of questions simultaneously have zero tolerance for inconsistent data definitions. Thi
Cursor has launched a feature enabling users to publish and share canvases—custom applications such as dashboards, reports, and internal tools—with their teams. This update extends Cursor's canvas fun
OpenAI introduced a new memory system for ChatGPT that allows users to review and control what the AI remembers across conversations through a memory summary feature. This update provides enhanced tra
This post appears to be a query from Swyx asking @_catwu for an update on a previously shared chart, specifically regarding developments or changes that occurred after the release of Opus 4.8. The pos
This entry documents a Wordle game result where the player solved puzzle #1,810 in 4 attempts, using the color-coded feedback system (black for incorrect letters, yellow for correct letters in wrong p
arXiv:2606.04916v1 Announce Type: new Abstract: Worker utility is not observed -- only its consequence is. Each gig transaction produces a single bit: accepted or rejected. We argue this structure poi
arXiv:2606.05159v1 Announce Type: new Abstract: Rigorous evaluation of learning-based robotic systems is an essential prerequisite for deployment. However, real-world test data is expensive to gather;
xAI announced a partnership with Vapi to integrate voice capabilities into xAI's AI systems and services. The collaboration aims to enhance conversational AI by leveraging Vapi's voice technology plat
arXiv:2606.04301v1 Announce Type: new Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target dataset
you guys know where this is going right wow this @reve 2.0 launch copy is supurb. 'it is now clear that the key to both controllable image generation and editing is not denser prompts, but a highly de
You know this is about Sam from line 1, even before you get to the picture. TL;DR: He’s going to keep bullshitting his way to IPO. This entire industry is based on an illusion. It deliberately mistake
arXiv:2512.17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpreta
arXiv:2606.04906v1 Announce Type: cross Abstract: Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detectio
arXiv:2606.04788v1 Announce Type: new Abstract: Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive m
arXiv:2603.09170v2 Announce Type: replace-cross Abstract: Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation d
arXiv:2606.05102v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict o
This benchmark demonstrates how integrating Alluxio with Ray Data achieves 20x faster training data read speeds for cross-region machine learning workloads on Anyscale's platform. The study shows perf
95% token reduction. 30x faster execution. 90%+ task completion. Today at #MSBuild, we announced a major shift to move reasoning upstream: Pinecone Nexus now integrates directly with @Microsoft OneLak
arXiv:2606.03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matte
arXiv:2606.03646v1 Announce Type: new Abstract: This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formul
This OpenAI publication outlines a proposed framework for establishing democratic oversight and governance structures for advanced AI systems, addressing how frontier AI development should be regulate
arXiv:2512.16882v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have brought substantial gains in the extrapolation capability in computational chemistry. How
arXiv:2606.03685v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) improves end-to-end classical planning in large language models (LLMs), but do these models also learn to represent and r
arXiv:2606.03156v1 Announce Type: new Abstract: We describe a versioned cross-domain dataset of 410,499 active tropical species (working snapshot 2026-04-20) spanning three applied subdomains -- tropi
arXiv:2511.13899v2 Announce Type: replace-cross Abstract: Low-rank recurrent neural networks (lrRNNs) are a class of models that uncover low-dimensional latent dynamics underlying neural population ac
arXiv:2606.03675v1 Announce Type: new Abstract: Methane is a potent greenhouse gas, and detecting leaks early via hyperspectral satellite imagery can help climate change mitigation efforts. Meanwhile,
arXiv:2606.03018v1 Announce Type: cross Abstract: Modeling interactions among multimodal, high-dimensional data is intrinsically challenging due to ultra-high dimensionality and complex dependence str
A few months ago, I found an anonymous sockpuppet account linked to the OpenAI/a16z super PAC. Now, @TaylorLorenz and I have uncovered two more — and they're even more brazen than the first. https://x
arXiv:2606.03471v1 Announce Type: new Abstract: This paper proposes, for the first time, a rigorous formal definition of the concept of Machine Theory of Mind, based on principles supported by evidenc