Another big langchain week
Another big langchain week 🚀langchain launches this week: all about open source models and memory! First: open source models. We partnered with @NVIDIAAI to launch a NemoClaw DeepAgents blueprint. Thi
Knowledge catalogue
Another big langchain week 🚀langchain launches this week: all about open source models and memory! First: open source models. We partnered with @NVIDIAAI to launch a NemoClaw DeepAgents blueprint. Thi
Call the homies, new @UnslothAI NVFP4 just dropped 🔥 We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We a
arXiv:2607.08436v1 Announce Type: cross Abstract: Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like objects, scene
arXiv:2607.07880v1 Announce Type: new Abstract: Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications
GROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARK New data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world professional tasks. On their GDPval+ benchmark (expert-created
'I think it's important for people to understand how code works.' Geoffrey Litt's Design Eng track keynote is live now: https://www.youtube.com/watch?v=WkBPX-oDMnA Thank you @NotionHQ for supporting h
arXiv:2509.26076v2 Announce Type: replace Abstract: As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on researc
OpenWiki general purpose memory is meant to be complementary to codex/claude code memory: it's proactive & ambient, meaning it'll automatically go out into your world (via connections like gmail, x, n
arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic
A Visual Introduction to Information Theory (bookmark it) Information Theory is such an beautiful and powerful subject. In the era of AI, it's worth spending time learning about it. Here is a highly-r
arXiv:2607.07690v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard
arXiv:2607.07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize co
ChatGPT Work == Claude Cowork ChatGPT Codex == Claude Code I kinda wish OpenAI created a single unified app surface for all work, coding or not, even though I get the UI/UX would be different Introduc
create a really interesting game called 'Don't Discuss Goblins' that should have some story elements but mostly be fun and fast moving - the goal is to avoid mentioning goblins, and this should be cha
arXiv:2508.17298v3 Announce Type: replace-cross Abstract: Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability t
Is Meta AI back? Haven't seen Mark post in three years here. Plus, the model is available via API. Not to mention the courage to announce it the same week as the long-awaited GPT-5.6. Great timing if
OpenWiki Brains 0.1.0 is officially released! We added a general-purpose memory brain to OpenWiki, in addition to the existing code brain. You can now use it to seamlessly setup a personal brain to tr
arXiv:2607.07498v1 Announce Type: cross Abstract: Testing is a major effort for the gaming industry, requiring a significant part of development budget and people power. We present a case study on a d
arXiv:2607.07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems
SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model t
// The Harness Effect // (bookmark it) Now more that ever pay very close attention to the orchestration harness and its effect on costs and performance. This study ran 22 evaluation tasks on six found
arXiv:2607.06875v1 Announce Type: new Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis
arXiv:2607.06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fac
arXiv:2511.06229v3 Announce Type: replace Abstract: This paper focuses on dynamic origin-destination matrix estimation (DODE), a crucial calibration process necessary for the effective application of
Grok 4.5 brings frontier performance across coding and knowledge work SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and O
Grok 4.5 context window will upgrade to 1M probably by next week SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and Opus 4
Jamf AI Governance is a capability within Jamf for Mac that enables IT and security teams to discover actively-used AI tools, enforce policy controls, and generate audit-ready reporting, providing com
Not another demo or benchmark. @ShoucongChen is a senior member of our technical staff. A real project, scoped at 1-month. Delivered in 4 days with GLM5.2 Fast. The best devs deserve >400 t/sec. Take
arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz
arXiv:2607.06558v1 Announce Type: new Abstract: Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demons
We are excited to launch 𝗥𝗲𝘀𝘁𝗮𝘁𝗲 𝗕𝗬𝗢𝗖 (Bring-your-own-Cloud) today. It is fully managed @restatedev, with all the features of Restate Cloud, but in your own account, in a dedicated VPC. Data never lea
We are hiring for @Harvey’s model training team. This team will help Harvey expand from the application layer into the model layer and from legal into high end knowledge work more broadly. We are hiri
We need AI model selection to be MUCH easier ASAP. I want proactive flags from my AI systems suggesting models. I want my AI harness to say 'hey allie, my girl, you keep asking for bar recommendations
Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. OpenAI is at 7.4%. For now, the frontier is American models,
arXiv:2504.11907v3 Announce Type: replace Abstract: Autonomous exploration of cluttered environments requires efficient exploration strategies that guarantee safety against potential collisions with u
arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.
arXiv:2509.08269v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly integrated with evolutionary computation to support optimization tasks. This survey primarily fo
arXiv:2607.04719v1 Announce Type: new Abstract: Aerial robots are increasingly moving from remote observation toward physical interaction with objects, surfaces, structures, loads, and surrounding flo
arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language M
arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con
arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders
arXiv:2607.02799v1 Announce Type: new Abstract: Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agen
arXiv:2607.03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenario
arXiv:2607.04653v1 Announce Type: new Abstract: While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional p
arXiv:2607.04103v1 Announce Type: cross Abstract: The release of SR 26-2 marks a significant modernization of U.S. model risk management by replacing SR 11-7 with a more risk-based and materiality-sen
arXiv:2607.03512v1 Announce Type: new Abstract: Existing classical control methods commonly require precise models and struggle to cope with model uncertainties and external disturbances, while end-to
arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid
arXiv:2510.08807v3 Announce Type: replace-cross Abstract: From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. Howev
arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them
arXiv:2607.02668v1 Announce Type: new Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generato
arXiv:2607.02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizati
arXiv:2409.16663v5 Announce Type: replace-cross Abstract: We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a ne
arXiv:2607.04972v1 Announce Type: cross Abstract: Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes, yet existin
arXiv:2607.04371v1 Announce Type: new Abstract: We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maxim
NEW AI paper worth bookmarking. This is something I called early, and this paper confirms it: verification has emerged as a new important scaling axis. Here is the simple explainer and what this paper
nvidia ai just handed dgx spark owners a real gift. nemotron labs 3 puzzle 75b a9b nvfp4 is basically built for this box. 75b total, 9.3b active, nvfp4, mamba plus moe, 256k context in the config, and
arXiv:2510.01764v3 Announce Type: replace Abstract: Reinforcement learning (RL) research requires diverse, challenging environments that are both tractable and scalable. While modern video games may o
arXiv:2607.03261v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in 3D spatial reasoning, spatial grounding, and fine-grained geometric underst
arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as pe
arXiv:2601.11049v2 Announce Type: replace-cross Abstract: We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions c