Visual Compositional Tuning
arXiv:2504.21850v3 Announce Type: replace Abstract: Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlo
Knowledge catalogue
arXiv:2504.21850v3 Announce Type: replace Abstract: Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlo
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. When Google opens its doors tomorrow for its annual developer
arXiv:2605.15959v1 Announce Type: cross Abstract: Physics-informed neural networks (PINNs) are powerful surrogates for differential equations but are notoriously difficult to train due to spectral bia
arXiv:2605.16035v1 Announce Type: cross Abstract: AI agents are increasingly deployed to act autonomously in the world, yet there is still no reliable way to trace a harmful agent back to the account
arXiv:2511.02342v3 Announce Type: replace Abstract: Aerial manipulation combines the maneuverability of multirotors with the dexterity of robotic arms to perform complex tasks in cluttered spaces. Yet
will post more about (final) day 2 of @aiDotEngineer Singapore soon but memorable highlights i wanted to remember is how @theesabina and @bytheophana (both from nyc) were genuinely shocked and on guar
FSD V14.3.3 is a banger Tesla FSD V14.3.3 just started rolling out, and it comes with the Spring Update! Actually Smart Summon’s top speed has also been increased to 8 mph (from 6 mph)! Downloading it
LlamaIndex held an offsite event in 2026, as announced by Jerry Liu, the project's founder, on social media. The post likely covered updates on the project's direction, team initiatives, or announceme
Recursive self-improvement is very reasonably the greatest near-term threat to democracy & peace out there New letter from 35 (!) members of Congress to the White House urging action post-Mythos. Most
The post argues that 'return on inference' (ROI) has become the primary metric for evaluating AI investments and business value over the next 18 months, shifting focus from traditional return on inves
symbolic tools like those listed below - rather than pure scaling -are likely driving most of the advances now. once you realize that you realize that hyperscaling is probably insane. @GaryMarcus IMHO
The salvation is Project Tapestry https://thealliance.ai/projects/tapestry I don't think people understand just how bad it will be if an American open source champion doesn't emerge soon and the big l
The search results don't contain specific information about the 'Wasteland Sweeper' post itself. Based on the context that it's posted to r/StableDiffusion, a community for sharing and discussing AI-g
Are your benchmarks actually measuring the capability you think they measure? New paper says they probably not. Coined the 'The Evaluation Trap', it provides a vocabulary for auditing whether your eva
Truly an all-star cast, on one of the most important questions in AI. Thrilled to see some many people finally willing to confront the hard questions of how we can move beyond LLMs, and into what worl
arXiv:2605.13850v1 Announce Type: new Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus o
arXiv:2512.01977v2 Announce Type: replace-cross Abstract: The global capacity for mineral processing must expand rapidly to meet the demand for critical minerals, which are essential for building the
also thought this was cool from the creative hackathon. extended the dashboard into a frontend dev agent. giving an agent a tight use case and surfacing it via a mini chat right next to the artifact i
arXiv:2605.14331v1 Announce Type: cross Abstract: Modern edge devices increasingly rely on neural networks for intelligent applications. However, conventional digital computing-based edge inference re
arXiv:2605.14393v1 Announce Type: new Abstract: We study analogical trajectory transfer, where the goal is to translate motion trajectories in one 3D environment to a semantically analogous location i
arXiv:2605.14518v1 Announce Type: new Abstract: Activation functions are central to deep networks, influencing non-linearity, feature learning, convergence, and robustness. This paper proposes the Ada
arXiv:2605.15198v1 Announce Type: cross Abstract: Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is
arXiv:2605.14231v1 Announce Type: cross Abstract: Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have b
b9163 is a llama.cpp release that adds a custom Unicode regex handler for Qwen3.5's tokenizer to prevent stack overflows on long inputs . The release also includes improvements to SYCL memory manageme
arXiv:2605.14106v1 Announce Type: new Abstract: We investigate whether behavior cloning is sufficient to produce active perception in a structured object-finding task. A low-cost robot arm equipped wi
arXiv:2508.02332v3 Announce Type: replace Abstract: The performance of Bayesian optimization (BO), a highly sample-efficient method for expensive black-box problems, is critically governed by the sele
arXiv:2605.13887v1 Announce Type: cross Abstract: Transformer-based Spiking Neural Networks (SNNs) integrate SNNs with global self-attention and have demonstrated impressive performance. However, exis
arXiv:2511.07308v2 Announce Type: replace Abstract: Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insigh
arXiv:2511.14751v2 Announce Type: replace Abstract: We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the
arXiv:2605.14795v1 Announce Type: new Abstract: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic sup
company that steals IP urges US government not to allow others to steal their IP Anthropic drops a paper on the US-China AI race They believe the US and its allies may be able to lock in a 12-24 month
arXiv:2605.14067v1 Announce Type: new Abstract: Financial distress prediction remains a significant challenge in enterprise risk analysis due to the highly imbalanced nature of real-world financial da
arXiv:2605.15062v1 Announce Type: new Abstract: Background. RGB-trained capsule-endoscopy classifiers underperform on small-vessel vascular findings by conflating hemoglobin contrast with bile and ill
arXiv:2605.14495v1 Announce Type: cross Abstract: Multimedia verification requires not only accurate conclusions but also transparent and contestable reasoning. We propose a contestable multi-agent fr
arXiv:2510.02952v3 Announce Type: replace Abstract: Inferring trajectories from longitudinal spatially-resolved omics data is fundamental to understanding the dynamics of structural and functional tis
arXiv:2509.01299v2 Announce Type: replace Abstract: Cross-domain few-shot segmentation (CD-FSS) aims to segment unseen categories with very limited samples while alleviating the negative effects of do
arXiv:2605.11611v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG)
arXiv:2405.07459v3 Announce Type: replace Abstract: Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS method
arXiv:2605.14220v1 Announce Type: cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match ex
arXiv:2507.04049v4 Announce Type: replace Abstract: Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstrations, often leading to conservative and homogeneo
arXiv:2605.14104v1 Announce Type: new Abstract: Inferring spatially resolved gene expression from histology images offers a cost-effective complement to spatial transcriptomics (ST). However, existing
arXiv:2502.00270v3 Announce Type: replace-cross Abstract: The performance of an LLM depends heavily on the relevance of its training data to the downstream evaluation task. However, in practice, the d
arXiv:2605.14014v1 Announce Type: cross Abstract: Internet of Things (IoT) systems continuously collect heterogeneous sensing signals from ubiquitous sensors to support intelligent applications such a
arXiv:2605.14742v1 Announce Type: new Abstract: Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing m
arXiv:2509.21261v3 Announce Type: replace Abstract: Micro-action Recognition is vital for psychological assessment and human-computer interaction. However, existing methods often fail in real-world sc
arXiv:2605.13853v1 Announce Type: cross Abstract: Facial editing is an important task with applications in entertainment, virtual reality, and digital avatars. Most existing approaches rely on generat
arXiv:2605.14868v1 Announce Type: new Abstract: Generating adversarial examples at scale is a core primitive for robustness evaluation, adversarial training, and red-teaming, yet even 'fast' attacks s
arXiv:2605.13974v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms t
arXiv:2605.14168v1 Announce Type: new Abstract: Learning of continuous exponential family distributions with unbounded support remains an important area of research for both theory and applications in
arXiv:2603.03577v2 Announce Type: replace Abstract: Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of
arXiv:2605.14445v1 Announce Type: new Abstract: Many real-world coding challenges are open-ended and admit no known optimal solution. Yet, recent progress in LLM coding has focused on well-defined tas
arXiv:2605.14501v1 Announce Type: cross Abstract: This paper proposes a fully dynamic Deep Reinforcement Learning (DRL) method for rebalancing dockless bike-sharing systems, overcoming the limitations
arXiv:2605.15018v1 Announce Type: cross Abstract: Shapley value and its priority-aware extensions are widely used for valuation in machine learning, but existing methods require pairwise priority to b
Had the honor of sharing the stage with the one and only @sydneyrunkle at Interrupt 2026 to talk about Deep Agents. Sign-up for the Managed Deep Agents waitlist today to get early access! https://www.
Honda previewed two new hybrid prototypes scheduled to launch by 2028 after absorbing a 9.2 billion EV-related loss, its largest in company history. The company indefinitely suspended its planned 15 b
arXiv:2602.01828v2 Announce Type: replace Abstract: Many complex networks exhibit hierarchical, tree-like structures, making hyperbolic space a natural candidate wherein to learn representations of th
A Reddit post showcasing a humorous AI-generated image created by asking ChatGPT to produce a still frame from a 1950s movie, where the comedic elements become increasingly apparent upon closer inspec
arXiv:2605.14774v1 Announce Type: new Abstract: In the world of AI and advanced technologies investigation aspects identification of a crime or criminal plays a major problem. In this research we focu
arXiv:2605.14239v1 Announce Type: new Abstract: Hyperspectral image (HSI) classification is challenging in complex scenes due to spectral ambiguity, spatial heterogeneity, and the strong coupling betw
arXiv:2605.14455v1 Announce Type: new Abstract: The Intelligence Impact Quotient (IIQ) is a composite metric intended to quantify the depth to which AI systems are integrated into organizational work