it’s funny how people here just make stuff up.
it’s funny how people here just make stuff up. @GaryMarcus Take away LLM & every single AI application today goes back to the stone age, driverless cars will immediately break down, all the apps will
Knowledge catalogue
it’s funny how people here just make stuff up. @GaryMarcus Take away LLM & every single AI application today goes back to the stone age, driverless cars will immediately break down, all the apps will
It’s the most wonderful time of the year 🎶 The 2026 AI Engineering Survey is live, this year in partnership with @NotionHQ and @vercel. If you’re an engineer working on AI products, we’d love you to f
arXiv:2605.16258v1 Announce Type: cross Abstract: Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most ex
arXiv:2506.23552v2 Announce Type: replace Abstract: The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS
arXiv:2605.16023v1 Announce Type: new Abstract: LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its
Judge Yvonne Gonzalez Rogers basically waved off Elon’s OpenAI appeal, saying the jury’s statute of limitations ruling is tough to overturn. An activist judge letting Sam Altman skate after he hijacke
just as i feared AI can be a powerful tool for opening people up to new perspectives, yet we find that people actually prefer to use “sycophantic” AI systems that reinforce their pre-existing beliefs.
arXiv:2605.15548v1 Announce Type: new Abstract: Traditional robotic hand metrics focus on static properties such as workspace, manipulability, and grasp stability. However, these metrics do not direct
arXiv:2605.15419v1 Announce Type: new Abstract: Flow matching trains a neural velocity field by regression against a target velocity associated with a prescribed probability path connecting a simple i
arXiv:2605.15769v1 Announce Type: cross Abstract: The co-optimization of a robot's body and brain presents a coupled challenge: the morphology constrains which control strategies are effective, while
LangSmith Engine automates the full agent fix loop — detecting failures, diagnosing causes and drafting PRs. But multi-model enterprises say a neutral observability layer still wins. http://venturebea
arXiv:2605.15496v1 Announce Type: cross Abstract: Neural distance fields offer a compact and continuous representation of 3D geometry, making them attractive for incremental LiDAR mapping. However, th
arXiv:2603.25099v2 Announce Type: replace-cross Abstract: We present a framework in which a large language model (LLM) acts as an online adaptive controller for SIMP topology optimization, replacing c
arXiv:2504.08300v5 Announce Type: replace-cross Abstract: Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Langua
arXiv:2512.19701v2 Announce Type: replace-cross Abstract: Accurate prediction of resource consumption and runtime for cloud workflow jobs is critical for scheduling efficiency, yet remains challenging
last week, a friend was telling me how running down a steep trail is great cognitive exercise due to the vast amount of quick processing you do from vision and tactile input to micro adjustments requi
In 2023, China's annual energy capacity additions exceeded Germany's total installed energy capacity, highlighting the scale of China's infrastructure expansion and renewable energy deployment. This c
arXiv:2605.15618v1 Announce Type: cross Abstract: Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score o
arXiv:2603.10881v2 Announce Type: replace Abstract: Electroencephalogram (EEG) classification plays a key role in medical diagnosis and brain-computer interfaces, but remains challenging due to low si
arXiv:2605.16234v1 Announce Type: cross Abstract: When researchers ask whether two transformer layers are 'equivalent' for compression, they often conflate distinct tests. Replacement asks whether one
arXiv:2605.15895v1 Announce Type: cross Abstract: Clinical application of high-resolution diffusion MRI is hindered by hardware limitations and prohibitive scan times, motivating computational super-r
arXiv:2605.15463v1 Announce Type: new Abstract: As machine learning models grow in complexity, they increasingly struggle with three conflicting demands: the need for high accuracy, the requirement fo
arXiv:2605.15582v1 Announce Type: new Abstract: Modern deep learning models for change detection (CD) often struggle to explicitly represent task-relevant semantic differences. This paper proposes the
arXiv:2605.15341v1 Announce Type: cross Abstract: LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let the
Composer 2.5 is a feature or update from Cursor, an AI-powered code editor, likely introducing improvements to its code generation and editing capabilities. The update probably includes enhancements t
arXiv:2605.16154v1 Announce Type: new Abstract: Reinforcement learning (RL) allows vision-language-action (VLA) policies to generalize beyond their training distribution by optimizing directly for tas
arXiv:2605.15760v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) optimization is most commonly performed using standard optimizers (Adam, SGD). While stable across diverse scenes, standard
arXiv:2605.15975v1 Announce Type: new Abstract: We tackle the challenge of building embodied AI agents that can reliably solve long-horizon planning problems. Imitation learning from demonstrations ha
arXiv:2605.15789v1 Announce Type: new Abstract: Uncertainty quantification is essential in safety-critical settings--from autonomous driving to aviation, finance, and health--where decisions must rely
arXiv:2605.15640v1 Announce Type: new Abstract: Multi-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existi
arXiv:2605.15713v1 Announce Type: cross Abstract: Legged manipulators extend robotic capabilities beyond static manipulation by integrating agile locomotion with versatile arm control. However, achiev
arXiv:2605.15535v1 Announce Type: new Abstract: Underwater salient object detection (USOD) has attracted increasing attention for underwater visual scene understanding and vision-guided robotic applic
arXiv:2504.09006v4 Announce Type: replace-cross Abstract: We initiate the study of structured Stackelberg games, a novel form of strategic interaction between a leader and a follower where contextual
arXiv:2605.15487v1 Announce Type: cross Abstract: Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from tw
arXiv:2605.15236v1 Announce Type: cross Abstract: With the coded caching, the server can use the information the users have cached to serve multiple users at a time by sending a single coded multi-cas
arXiv:2605.16043v1 Announce Type: cross Abstract: Deformable Linear Objects (DLOs) such as ropes and cables are widely encountered in both household and industrial applications, yet remain challenging
arXiv:2604.02812v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have recently demonstrated strong capabilities in mapping multimodal observations to robot behaviors. However, most cu
arXiv:2602.04909v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fix
arXiv:2605.15504v1 Announce Type: cross Abstract: Financial, social, and political factors often prevent the interests of the owners of ML systems and services and their users from being perfectly ali
arXiv:2605.15886v1 Announce Type: new Abstract: This paper introduces a dataset of interlinked multimodal political communications from the Russian government, addressing persistent deficiencies in th
Literary journals are now publishing, and awarding prizes to, AI written stories. Surprised this made it into Granta! ‘The Serpent in the Grove’ by Jamir Nazir is a story set in rural Trinidad about a
llama.cpp adds MTP for the Qwen3.6 family This is a significant milestone for the local AI ecosystem. The performance jump with these changes is massive and elevates local inference on commodity hardw
llama.cpp with MTP support makes local models fast enough to use as daily drivers 🚀 Qwen3.6-27B dense generation (on A10G): From 25 tok/s → 45 tok/s (+78%). Two flags on llama-server: --spec-type draf
LLM apps fail in ways normal logs can’t explain. Same input, different outputs. Subtle drift. Silent hallucinations. That’s the problem space LangSmith by @LangChain was built for: observability + eva
arXiv:2511.19931v2 Announce Type: replace-cross Abstract: Cross-domain Sequential Recommendation (CDSR) has been proposed to enrich user-item interactions by incorporating information from various dom
arXiv:2605.15916v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has emerged as an critical technique for adapting large-scale foundation models across natural language process
arXiv:2509.21663v2 Announce Type: replace-cross Abstract: Neurosymbolic integration (NeSy) blends neural-network learning with symbolic reasoning. The field can be split between methods injecting hand
arXiv:2605.15242v1 Announce Type: new Abstract: The reliability of Healthcare Information Systems (HIS) is frequently compromised by human-induced data entry errors, which existing statistical anomaly
arXiv:2602.23409v2 Announce Type: replace-cross Abstract: Angle-encoded variational quantum circuits admit a truncated Fourier series representation of their output, but approximating functions with m
arXiv:2605.16143v1 Announce Type: new Abstract: Large language model based agents often fail in unfamiliar environments due to premature exploitation: a tendency to act on prior knowledge before acqui
arXiv:2605.16048v1 Announce Type: cross Abstract: State Space Models (SSMs) are inherently recurrent along the sequence dimension, yet depth-recurrence - reusing the same block repeatedly across layer
The post discusses positive features and improvements included in version 0.6 of DeepAgents, with Harrison Chase praising a write-up by Sydney that explains these updates. This likely covers new capab
arXiv:2605.15393v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed to perform tasks with minimal human oversight, it is crucial that these models operate robustl
arXiv:2605.15621v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, e
arXiv:2509.05030v2 Announce Type: replace Abstract: To enable large-scale reuse of real-world 3D assets, where garments and characters rarely share skeletons, templates, or dense correspondences, we p
arXiv:2605.16179v1 Announce Type: new Abstract: Agricultural landscape segmentation in the Global South is challenging as it is characterized by fragmented plots, high intra-class variance, and a scar
arXiv:2605.15416v1 Announce Type: cross Abstract: Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relyi
arXiv:2605.15806v1 Announce Type: new Abstract: Neural operators excel as deterministic surrogates, but inevitably collapse to the conditional mean when applied to stochastic PDEs, discarding the vari
arXiv:2605.15231v1 Announce Type: cross Abstract: Nonlinear finite element crash simulations are accurate but computationally expensive, limiting their use in iterative design optimisation. Machine-le
@masondrxy and @Vtrivedy10 are doing an awesome push on harness profiles @huntlovell doing cool stuff w code interpreter / REPL @bromann leading the charge on streaming our OSS team is stacked! lots o