Caracal: Causal Architecture via Spectral Mixing
arXiv:2605.00292v1 Announce Type: new Abstract: The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. T
Knowledge catalogue
arXiv:2605.00292v1 Announce Type: new Abstract: The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. T
claudely is a tool that enables users to run Claude Code against local LLM providers such as LM Studio, Ollama, or llama.cpp while preserving their existing Claude configuration. The tool allows devel
arXiv:2605.00630v1 Announce Type: new Abstract: The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video
arXiv:2605.00222v1 Announce Type: new Abstract: Chemical reaction datasets such as USPTO suffer from substantial incompleteness, frequently missing byproducts, co-reactants, and stoichiometric coeffic
arXiv:2509.06864v2 Announce Type: replace Abstract: This paper introduces PyFair, a formal framework for evaluating and verifying individual fairness of Deep Neural Networks (DNNs). By adapting the co
arXiv:2605.00513v1 Announce Type: new Abstract: Understanding how people argue across ideological divides online is important for studying political polarization, misinformation, and content moderatio
arXiv:2605.00350v1 Announce Type: new Abstract: ``How long can I live and remain free of cancer?'' is often the first question a patient asks after receiving a cancer diagnosis and treatment. Accurate
deepagents-cli is quietly becoming the best place to start coding with open weight models. we've been investing heavily in making it a harness that's truly model-agnostic, without compromising perform
arXiv:2509.21864v2 Announce Type: replace Abstract: The wide availability and low usability barrier of modern image generation models has triggered the reasonable fear of criminal misconduct and negat
Deepseek V4 works more thoroughly than other open source models: It writes its own tests and performs extensive validation. This leads to better performance, but also cases of the model being overconf
arXiv:2510.00233v2 Announce Type: replace Abstract: Scientific machine learning has enabled the extraction of physical insights and data-driven modeling of high-dimensional spatiotemporal data, yet ac
arXiv:2605.00066v1 Announce Type: new Abstract: Open-loop evaluation offers fast, reproducible assessment of autonomous driving planners, but its ability to predict real closed-loop driving performanc
This documentation page describes how to integrate Ollama with Claude Desktop, enabling users to run local language models through the Anthropic Claude interface. The integration allows Claude Desktop
arXiv:2602.18757v2 Announce Type: replace Abstract: Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecti
arXiv:2410.04299v2 Announce Type: replace Abstract: Incorporating a priori physics knowledge into machine learning leads to more robust and interpretable algorithms. In this work, we combine deep lear
arXiv:2605.00296v1 Announce Type: new Abstract: Plant phenology-the study of recurrent life cycle events-is essential for understanding ecosystem dynamics and their responses to climate change impacts
arXiv:2605.00677v1 Announce Type: new Abstract: While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results ste
arXiv:2504.05679v2 Announce Type: replace Abstract: Small unmanned aerial vehicle (UAV)-based visual inspections are a more efficient alternative to manual methods for examining civil structural defec
Christopher Mims / Wall Street Journal: Ex-iRobot CEO Colin Angle launches Familiar Machines & Magic and unveils Familiar, a dog-like, “emotionally intelligent” robot that reacts to owner's feelings —
arXiv:2507.14201v3 Announce Type: replace-cross Abstract: We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions
arXiv:2504.10368v4 Announce Type: replace Abstract: This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal tok
arXiv:2605.00605v1 Announce Type: new Abstract: Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-pos
arXiv:2605.00011v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative intelligence across decentralized data source devices in a privacy-preserving way. While substantial resea
arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal
In the era of AI agents, the distance between a big idea and a working application has never been shorter. As we lean more heavily on agents to help us build applications, a critical question remains:
arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim
arXiv:2605.00420v1 Announce Type: cross Abstract: Evaluating the true forecasting ability of AI agents requires environments resistant to overfitting, free from centralized trust, and grounded in ince
arXiv:2605.00358v1 Announce Type: new Abstract: LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the
arXiv:2605.00147v1 Announce Type: new Abstract: On-orbit inspection imagery is crucial as it enables characterization of non-cooperative resident space objects, providing the geometry and structural c
arXiv:2605.00645v1 Announce Type: new Abstract: Clinical time-series forecasting is increasingly studied for decision support, yet standard aggregate metrics can obscure whether a model is actually us
SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture, marking a fundamental paradigm shift from mo
Granite 4.1 3B SVG Pelican Gallery IBM released their Granite 4.1 family of LLMs a few days ago. They're Apache 2.0 licensed and come in 3B, 8B and 30B sizes. Granite 4.1 LLMs: How They’re Built by Gr
Grok 4.3 just built this entire game with just a single prompt It has the fastest output token speed and outranks Claude Sonnet 4.6 Max on Artificial Analysis I built this using the xAI API in Kilo Co
arXiv:2605.00281v1 Announce Type: new Abstract: We study high-probability (HP) convergence guarantees in decentralized stochastic optimization, where multiple agents collaborate to jointly train a mod
arXiv:2605.00113v1 Announce Type: new Abstract: We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe
This post documents a 10-day project to build a local AI system using open-source tools and models, specifically combining Ollama (a local LLM framework), Gemma 4 (a language model), Claude Cowork 3P,
OpenAI describes its technical approach to delivering real-time voice AI services with minimal latency across large user bases, likely covering infrastructure optimization, model serving strategies, a
arXiv:2507.01955v3 Announce Type: replace Abstract: Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond que
arXiv:2605.00351v1 Announce Type: new Abstract: Root cause localization in cloud native microservice systems requires modeling complex service dependencies, irregular temporal dynamics, and heterogene
I had Claude Code for web build me this WebAssembly playground for trying out the new Redis array commands https://tools.simonwillison.net/redis-array More notes here: https://simonwillison.net/2026/M
I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for peopl
arXiv:2508.07630v2 Announce Type: replace Abstract: We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task
Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task of training a tiny MoE with all the architectural advancements
arXiv:2605.00184v1 Announce Type: new Abstract: With the growing integration of human-computer interaction into everyday life, advances in machine learning have enabled systems to better perceive and
arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta
arXiv:2605.00267v1 Announce Type: new Abstract: As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this
arXiv:2509.21514v3 Announce Type: replace-cross Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment r
arXiv:2605.00448v1 Announce Type: new Abstract: The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetri
arXiv:2510.19897v2 Announce Type: replace Abstract: We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without p
arXiv:2605.00051v1 Announce Type: new Abstract: Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions
arXiv:2412.00452v2 Announce Type: replace-cross Abstract: Conventioanl federated learning (FL) heavily depends on high-quality labels, which are often impractical in the real world, leading to the fed
arXiv:2505.22003v2 Announce Type: replace Abstract: In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full adv
arXiv:2605.00505v1 Announce Type: cross Abstract: Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented g
This podcast episode from Last Week in AI discusses recent developments in large language models, including updates on GPT 5.5 and DeepSeek V4, while also covering concerns about potential sabotage or
arXiv:2605.00398v1 Announce Type: new Abstract: Causal graph discovery for space-time systems is challenging in high-dimensional gridded data, which often has many more grid cells than temporal observ
arXiv:2605.00789v1 Announce Type: new Abstract: Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency
I cannot provide a summary for this entry as the post content appears incomplete or corrupted in the provided information. The title is cut off mid-sentence and doesn't clearly convey the full topic.
arXiv:2510.17281v5 Announce Type: replace Abstract: Scaling up data, parameters, and test-time computation has been the mainstream methods to improve LLM systems (LLMsys), but their upper bounds are a
arXiv:2510.23557v2 Announce Type: replace-cross Abstract: Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of h
arXiv:2505.23723v2 Announce Type: replace Abstract: The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering.