Microsoft’s first advanced reasoning AI is here
Microsoft announced a bunch of new in-house AI models at Build 2026, including a new 'flagship' model: MAI-Thinking-1. It's an ambitious step into model development for Microsoft, which introduced its
Knowledge catalogue
Microsoft announced a bunch of new in-house AI models at Build 2026, including a new 'flagship' model: MAI-Thinking-1. It's an ambitious step into model development for Microsoft, which introduced its
arXiv:2606.02031v1 Announce Type: cross Abstract: Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despit
Remind me, @Elonmusk, was GPT-5 really smarter than the smartest humans? @viktaur27 @Teslarati The rate of improvement from original GPT to GPT-3 is impressive. If this rate of improvement continues,
Harshita Mary Varghese / Reuters: Sensor Tower: ChatGPT has become the fastest app to hit 1B global MAUs by far; ChatGPT's MAUs are up 62% YoY in Q2 to date, Claude's MAUs are up 640% YoY to 56M — Ope
arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif
Device intelligence company FingerprintJS Inc. today launched a preview of two products built to identify traffic from artificial intelligence assistants, addressing a detection gap that has opened as
arXiv:2605.30813v1 Announce Type: new Abstract: We propose a novel algorithm for incremental Byte Pair Encoding (BPE) tokenization. The algorithm processes each input byte in worst-case O(log^2 t) tim
One of the new, buzzy jobs in Silicon Valley is the AI Forward Deployed Engineer (FDE), an engineer who is embedded within a client organization to help customize solutions, such as building and tunin
We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se
weird the way this tweet was getting a ton of traffic and then just stopped. 🤷♂️ I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthro
arXiv:2605.29379v1 Announce Type: new Abstract: We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class whi
arXiv:2605.29927v1 Announce Type: cross Abstract: Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints
arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration
arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red
Today I learned that @Grimezsz has more courage in her pinky than Roon does in his entire cowardly body. Roon made excuses; Grimes defended her views calmly and respectfully, like grownups should. Ope
crazy that this was announced on the day tokenmaxxing died. Anthropic raised 65 billion in a funding round that valued the artificial intelligence company at 965 billion including the new investment,
arXiv:2605.27766v1 Announce Type: new Abstract: LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongsi
arXiv:2501.01669v4 Announce Type: replace Abstract: Inverse reinforcement learning (IRL) has progressed significantly toward accurately learning the underlying rewards in both discrete and continuous
arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r
arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or
Artificial intelligence training data provider Human Archive Inc. today announced that it has raised 8.2 million in funding. Wing Venture Capital, NVP Capital, Y Combinator headlined the consortium th
arXiv:2605.27072v1 Announce Type: cross Abstract: We present E3, an automated review assistant that augments reviewers and engineering teams by identifying decision-relevant technical concerns in rese
I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from production traces to self-improve via detailed tracing tightly
“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa
U.S. law enforcement agencies are monitoring growing backlash to AI and have begun classifying anti-technology sentiment as an extremism threat, with unpublished reports from the Department of Homelan
arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL
arXiv:2605.24053v1 Announce Type: new Abstract: Large Language Models (LLMs) are predominantly governed by probabilistic frameworks in which the sum of outcome probabilities is constrained to unity. T
arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring
arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align
arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h
Nano Banana Pro vs. GPT Image 2 Nano Banana Pro wins on photorealism, 4K output, 14 reference image slots for product/scene consistency, and live web search for real-world accuracy. GPT Image 2 wins o
arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC
New York Times: A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies — The government's A.I. Se
An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3D Design, ahead of models 10X more expensive like Opus 4.7 b
arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds
arXiv:2605.23415v1 Announce Type: cross Abstract: Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant M
arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis
I once asked a bunch of e/accs how much damage was an acceptable risk relative to take (any) safety precautions. none answered. we may be about to find out. One implication of the below is that we rea
it ain’t just me who sees the emperor has no clothes The AI bubble math doesn't add up. Anthropic spends 3 to make 1 and that’s before you include any and all other costs like staff or electricity. Mi
arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr
arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema
I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at stake. We don’t know for example • How man
arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal
can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever
proof too complicated, Claude help ELI5 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematician
🚀Qwen3.7-Max just landed at 56.6 on the Artificial Analysis Intelligence Index — a solid 4.8pt jump over Qwen3.6-Max-Preview. @ArtificialAnlys ⚡️Sharper sci reasoning, stronger agentic chops, better c
It’s a really special time to be alive…some thoughts from training this model 🧵 Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in
Once AI starts making solving open problems in novel ways it won’t stop. We are entering the final stage of human solutions to open problems like this. Feels weird, doesn’t it? Today, we share a break
arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in
the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled changed or how the model does on anything else. (or what it costs
three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating everyone in achieving their goals today it was great to announce
arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th
arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i
arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a
Artificial intelligence developer Decart.ai Inc. today announced that it has raised 300 million in funding at a nearly 4 billion valuation. Radical Ventures led the round with participation from Nvidi
arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti
arXiv:2605.17554v1 Announce Type: new Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand. They are being de
arXiv:2511.21654v2 Announce Type: replace Abstract: We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in w
arXiv:2605.17246v1 Announce Type: cross Abstract: We introduce fidelity probes: natural-language questions generated from a reference artifact with code-derived ground-truth answers, answered from a c
I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning talk at PyCon US 2