DeepLook: Deeper Thinking with Lookahead
arXiv:2607.22602v1 Announce Type: new Abstract: Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reaso
Knowledge catalogue
arXiv:2607.22602v1 Announce Type: new Abstract: Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reaso
arXiv:2607.24165v1 Announce Type: cross Abstract: Finding a long document relevant to a multi-part request is not the same as establishing that it contains every requested piece of evidence. We study
arXiv:2607.23845v1 Announce Type: new Abstract: Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient sign
arXiv:2607.23614v1 Announce Type: cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates
arXiv:2607.24560v1 Announce Type: cross Abstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer
arXiv:2607.23899v1 Announce Type: cross Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot
Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Kimi K3 I would not miss this one! Kimi K3 has everyone’s atte
arXiv:2510.03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastru
arXiv:2607.24459v1 Announce Type: new Abstract: Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subse
arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates el
arXiv:2607.22959v1 Announce Type: cross Abstract: AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control
arXiv:2607.24562v1 Announce Type: new Abstract: Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives r
How OpenAI hacked HuggingFace. What we know. Hugging Face proved that open platforms and open models can still win those battles when the alternative is locked-down systems that refuse to assist their
Most quantization works like this: pick a bit depth, apply it everywhere, maybe let imatrix take a rough guess at what matters, ship it. Most don't check which specific weight groups can take a hit an
arXiv:2607.23153v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their correspondi
arXiv:2607.23242v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly
Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Jue Wang and Zain Hasan from Together AI for a deep dive into
arXiv:2607.23515v1 Announce Type: new Abstract: Long-horizon manipulation tasks pose significant challenges for reinforcement learning due to sparse reward signals and long horizons. Automatic curricu
arXiv:2607.24101v1 Announce Type: cross Abstract: Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing evaluations
arXiv:2607.22832v1 Announce Type: new Abstract: Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed. Representing policies as executabl
Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a
arXiv:2602.05547v2 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment
arXiv:2607.18368v2 Announce Type: replace Abstract: Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-pol
arXiv:2607.23023v1 Announce Type: new Abstract: Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing
arXiv:2607.23310v1 Announce Type: cross Abstract: We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must be assigned
arXiv:2607.23617v1 Announce Type: new Abstract: Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulti
arXiv:2607.23794v1 Announce Type: cross Abstract: Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morpholog
arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact wit
arXiv:2607.23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity i
really grateful to @realbasilchatha for helping me curate a survey of the entire field of Forward Deployed Engineering in one track! From FDE 101 by @zkevinbai (Anthropic, Palantir, Rippling Founding
arXiv:2607.24352v1 Announce Type: new Abstract: The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables
arXiv:2603.03536v2 Announce Type: replace-cross Abstract: Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an un
arXiv:2607.23840v1 Announce Type: new Abstract: Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech.
arXiv:2607.23991v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However,
arXiv:2607.24304v1 Announce Type: cross Abstract: We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and soc
arXiv:2607.22575v1 Announce Type: new Abstract: Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this abili
arXiv:2607.23108v1 Announce Type: new Abstract: While scaling laws for imitation learning have primarily focused on generalization in open-world settings, the relationship between data and precision i
arXiv:2607.24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development
arXiv:2509.06108v3 Announce Type: replace-cross Abstract: Graph drawing concerns the algorithmic visualization of graphs. A good drawing of a graph is easy to read and facilitates solving tasks on the
Very interesting paper on LLM reasoning. They find that frontier models can exhibit invisible reasoning by leveraging semantically irrelevant filler tokens. In other words, invisible reasoning can ser
We recently joined NVIDIA, Microsoft, and others in supporting open-weight AI. Today, we’re putting that belief into the product with model choice on Replit, starting with Kimi K3. The future is the r
arXiv:2607.23469v1 Announce Type: cross Abstract: Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled par
arXiv:2607.21960v1 Announce Type: new Abstract: This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environme
arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmen
arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Exist
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essentia
I've been seeing a lot of news about the latest gemma 4 and qwen 3.6 being really good and the current go-to models but those are out of reach for my GPU at the moment. With 4GB VRAM and 40 GB RAM, I
arXiv:2607.10428v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: person
Boston-based startup CollectivIQ Inc. is targeting runaway artificial intelligence costs with a platform intended to give businesses control over how models are used. The company describes its product
Fresh off the launch of Opus 5, Claude Code creator Boris Cherny (@bcherny) joins YC's @sdianahu at Startup School 2026 to talk about what the newest models can do, how Claude Code came to be, and wha
Great technical paper from Harvard and MIT. It's on role drift in compound LLM systems. (bookmark it) End-to-end RL improves the accuracy of a multi-module LLM pipeline without constraining how the mo
tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend &
arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchm
The National Oceanic and Atmospheric Administration (NOAA) is embarking on a transformative journey to redefine how we understand and predict patterns in the Earth’s atmosphere that affect the weather
NVIDIA Ising Calibration 1.5 is a 31‑billion‑parameter vision‑language model that diagnoses and tunes quantum processors, delivering state‑of‑the‑art zero‑shot and in‑context learning on the QCalEval
One of the least explored questions in AI is how to build rich open-ended environments for learning. Benchmarks evaluate intelligence, environments shape it. iLands is interesting because it doesn't j
arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet
arXiv:2607.21612v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following,
Proud to be an inaugural partner alongside @NVIDIA and other industry leaders in the Open Secure AI Alliance. We look forward to accelerating the adoption and trust of open models and harnesses throug
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: