Scaling Inherently Interpretable Language Models
arXiv:2608.07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods w
Knowledge catalogue
arXiv:2608.07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods w
arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It co
arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en
arXiv:2608.07923v1 Announce Type: new Abstract: Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training
arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multi
arXiv:2608.08106v1 Announce Type: new Abstract: Synthetic Aperture Sonar (SAS) is core for wide-area detection of small underwater targets. However, large-scale, high-quality SAS datasets are scarce,
arXiv:2608.08594v1 Announce Type: new Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong
arXiv:2608.09568v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing
arXiv:2608.07531v1 Announce Type: cross Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing ext
arXiv:2608.08075v1 Announce Type: cross Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archi
arXiv:2608.07924v1 Announce Type: cross Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based dr
arXiv:2608.09763v1 Announce Type: new Abstract: Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it acros
arXiv:2608.09049v1 Announce Type: new Abstract: Mobile app reviews are a rich, continuously renewing source of how users experience privacy and security, yet existing taxonomies of these concerns are
arXiv:2410.01227v2 Announce Type: replace-cross Abstract: In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of their
arXiv:2608.07916v1 Announce Type: new Abstract: Image demosaicing reconstructs a full-color image from incomplete color measurements produced by a sensor covered with a color filter array (CFA). Most
arXiv:2608.09449v1 Announce Type: new Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefor
arXiv:2510.15968v2 Announce Type: replace-cross Abstract: Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDE-solving-based methods, while accurate,
arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that dem
arXiv:2608.07553v1 Announce Type: new Abstract: Robotic manipulation often requires object specific programming, manual data annotation, or calibrated perception pipelines, which limits rapid deployme
arXiv:2608.08685v1 Announce Type: new Abstract: Reliable semi-dense matching is essential for modern geometric vision systems. Designed under a coarse-to-fine paradigm, it achieves an optimal balance
arXiv:2607.28069v2 Announce Type: replace Abstract: Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix
arXiv:2608.07835v1 Announce Type: new Abstract: Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes su
arXiv:2608.08839v1 Announce Type: cross Abstract: World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actio
arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many ex
arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanis
arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per
arXiv:2608.08884v1 Announce Type: new Abstract: As collaborative robots have entered domains such as manufacturing, agriculture, and healthcare, programming or adapting robot behavior typically requir
arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly
arXiv:2608.09400v1 Announce Type: cross Abstract: Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Po
arXiv:2608.09201v1 Announce Type: new Abstract: Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space.
arXiv:2608.09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Tra
arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks
Bloomberg: Singapore raises its 2026 GDP growth forecast to 4.5%-5.5% from 2%-4%, citing stronger-than-expected global AI investment and improved external demand — Singapore upgraded its 2026 economic
arXiv:2512.09062v2 Announce Type: replace Abstract: Accurate 3D scene interpretation in active construction sites is essential for progress monitoring, safety assessment, and digital twin development.
arXiv:2608.09185v1 Announce Type: cross Abstract: Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process involving c
arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s
arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app
arXiv:2608.09410v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observat
arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr
arXiv:2608.08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security propertie
arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Age
arXiv:2608.08840v1 Announce Type: new Abstract: Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific bod
arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every
arXiv:2602.17632v3 Announce Type: replace-cross Abstract: Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-b
arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains thr
Data interoperability is quickly becoming a practical requirement for companies trying to move artificial intelligence into production. Picking the right model or adding computing capacity is only par
arXiv:2601.19180v2 Announce Type: replace-cross Abstract: Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approac
arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp
arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat
arXiv:2608.09271v1 Announce Type: cross Abstract: Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group n
arXiv:2608.08984v1 Announce Type: new Abstract: Model selection for imbalanced binary classification often uses the Matthews correlation coefficient (MCC), but thresholding makes validation rankings t
arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity
Stephanie Palazzolo / The Information: Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation —
Wall Street Journal: Sources: Anthropic is courting investors for what could be the biggest IPO ever, touting rapid growth and plans to address mounting public backlash against AI — Company is racing
Business Insider: Sources: former OpenAI Chief Product Officer Kevin Weil is aiming to raise 150M for a new AI science startup, seeking a valuation of at least 750M — Kevin Weil, the former chief prod
The Information: Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models — Nvidia is doubling down
arXiv:2603.22213v2 Announce Type: replace-cross Abstract: While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data
arXiv:2608.09887v1 Announce Type: new Abstract: Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Exi
Zac Hall / 9to5Mac: SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users — SpaceXAI and Cursor
arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c