Critique of Agent Model
arXiv:2606.23991v1 Announce Type: new Abstract: What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and
Knowledge catalogue
arXiv:2606.23991v1 Announce Type: new Abstract: What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and
arXiv:2606.24586v1 Announce Type: new Abstract: Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimizat
Fugu Ultra by @SakanaAILabs is live on OpenRouter! Excited to see more multi-model systems pushing the frontier. Fugu-Ultra is now live on @OpenRouter! ⚡ We share a core vision with the OpenRouter tea
arXiv:2508.16420v3 Announce Type: replace Abstract: Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an un
arXiv:2407.06136v4 Announce Type: replace Abstract: Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes from limited examples while preserving knowledge of previously
arXiv:2606.24644v1 Announce Type: new Abstract: Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper st
arXiv:2606.23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a g
arXiv:2606.24263v1 Announce Type: new Abstract: Microwave satellite imagery plays a crucial role in monitoring tropical cyclone precipitation and intensity worldwide, but suffers from long revisit tim
arXiv:2606.24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, most existing methods for auditing differential pr
arXiv:2606.24101v1 Announce Type: cross Abstract: Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer
arXiv:2606.24388v1 Announce Type: new Abstract: We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is designed to be di
arXiv:2606.24251v1 Announce Type: new Abstract: Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are incre
arXiv:2606.24610v1 Announce Type: new Abstract: The evaluation of cultural grounding context becomes complex when multiple cultures convey the same moral lesson. This challenge is particularly relevan
arXiv:2606.24842v1 Announce Type: new Abstract: In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently,
arXiv:2606.23473v1 Announce Type: new Abstract: Universal segmentation models exhibit significant potential for diverse tasks involving different imaging modalities and segmentation objectives. Task-I
arXiv:2606.21212v1 Announce Type: new Abstract: Causal discovery is critical for understanding complex data-generating mechanisms, yet traditional algorithms often struggle with highly non-linear and
arXiv:2606.23090v1 Announce Type: new Abstract: Cross-embodiment data have become central to training robotic foundation models. To leverage such heterogeneous data, we focus on flow-based object mani
arXiv:2606.22378v1 Announce Type: new Abstract: Event cameras enable high-frequency visual perception with microsecond latency, offering advantages for dynamic scenes. However, event-based small objec
arXiv:2602.20070v3 Announce Type: replace Abstract: We develop a kernel method for generative modeling within the stochastic interpolant framework, replacing neural network training with linear system
arXiv:2606.23406v1 Announce Type: new Abstract: We present HyperQuant (Hadamard, optimallY Packing, Entropy Rice-coding), a unified post-training quantization pipeline for the weights and the KV cache
arXiv:2606.20627v1 Announce Type: cross Abstract: Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide p
arXiv:2411.19468v4 Announce Type: replace Abstract: The random feature (RF) method is a powerful kernel approximation technique, but it typically uses fixed activation functions, limiting its adaptabi
arXiv:2606.21292v1 Announce Type: new Abstract: We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic repr
arXiv:2508.13744v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handli
arXiv:2606.22030v1 Announce Type: cross Abstract: We present Nous, a novel agent memory architecture grounded in the principle that knowledge is prediction, not storage. Rather than persisting facts a
arXiv:2606.23668v1 Announce Type: new Abstract: Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a
arXiv:2606.20764v1 Announce Type: new Abstract: Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, rea
arXiv:2606.22540v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execut
arXiv:2606.23524v1 Announce Type: new Abstract: End-to-end OCR increasingly relies on autoregressive sequence models, where the quadratic cost of Transformer attention limits efficient transcription o
arXiv:2606.23567v1 Announce Type: new Abstract: Masked diffusion language models decode by iteratively unmasking tokens, where the unmasking order defines an 'order of thought' that strongly influence
arXiv:2606.23041v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the
arXiv:2606.20197v2 Announce Type: replace Abstract: Actor-Critic Model Predictive Control (MPC) effectively addresses complex, non-convex control problems, but guaranteeing the closed-loop stability o
arXiv:2606.23166v1 Announce Type: new Abstract: There has been rapid progress in generative artificial intelligence (AI) models for inorganic crystal design, which can efficiently generate large numbe
arXiv:2606.22858v1 Announce Type: new Abstract: As machine learning models grow more influential and opaque, algorithmic fairness and explainability are critical for ensuring accountability. However,
arXiv:2606.21295v1 Announce Type: new Abstract: Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wise dynamics, w
arXiv:2606.20670v1 Announce Type: new Abstract: Wireless foundation models offer a path toward reusable channel state information (CSI) intelligence for sixth-generation (6G) systems. However, existin
arXiv:2606.16917v3 Announce Type: replace Abstract: We present Unified Motion-Action (UMA) Model, an approach that uses 3D object motion trajectories as a shared interface to bridge visuomotor control
arXiv:2606.22136v1 Announce Type: new Abstract: Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and s
What should a world model for agile quadrotor control actually provide? 📄 Arxiv: https://arxiv.org/pdf/2606.23444 🌐 Project: https://pratyaksh10.github.io/skyjepa-project-page/ 💻 Code: https://github.
arXiv:2606.21641v1 Announce Type: new Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that 'warm-start' search from prior knowledge, proposing s
Guess which is Fugu Ultra? This is how recent models compare when generating endless procedural terrain (using Three.js). All of these are one-shotted! Just wild! Trying a few more examples. Will shar
This article discusses using local language models to automatically triage pull requests in the OpenClaw repository at no cost, eliminating the need for expensive API-based solutions. The post likely
arXiv:2606.11446v1 Announce Type: new Abstract: This research introduces a framework for incorporating Concept Bottleneck Models (CBMs) into 3D generative architectures to address the inherent 'semant
arXiv:2606.11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based s
arXiv:2606.12218v1 Announce Type: cross Abstract: Understanding spatial distribution of fallow land is important for optimizing the food-water (FW) nexus, given fallowing's role in crop rotation and w
arXiv:2606.12113v1 Announce Type: cross Abstract: Transformer-based language models for SMILES strings suffer from a locality gap: standard character-level tokenization fragments chemically meaningful
arXiv:2606.11893v1 Announce Type: cross Abstract: The correspondence between large language models (LLMs) and the neural mechanisms underlying human higher-order cognition remains insufficiently chara
arXiv:2606.11206v1 Announce Type: new Abstract: Supervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and lim
arXiv:2606.12245v1 Announce Type: cross Abstract: Cold-start item recommendation remains a persistent challenge in real-world systems due to the absence of interaction histories. While prior models at
arXiv:2602.08735v3 Announce Type: replace Abstract: While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, whic
arXiv:2606.12195v1 Announce Type: new Abstract: Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts large
arXiv:2606.12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle
arXiv:2605.06485v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have transformed artificial intelligence, but their computational requirements remain prohibitive for most users.
arXiv:2606.11278v1 Announce Type: new Abstract: In this paper, we introduce the Soft Lamprey-Inspired Dual Environment Robot (SLIDER) and a proper modeling and optimization procedure employed to desig
arXiv:2606.11794v1 Announce Type: cross Abstract: Neurodegenerative diseases such as Alzheimer's disease (AD) require accurate and scalable tools for assessing disease severity, yet current clinical s
arXiv:2606.11792v1 Announce Type: cross Abstract: Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated respo
arXiv:2410.12327v2 Announce Type: replace Abstract: Large language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting re
arXiv:2606.11399v1 Announce Type: new Abstract: Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cul
arXiv:2606.12316v1 Announce Type: new Abstract: ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many
arXiv:2606.12006v1 Announce Type: cross Abstract: Predicting time-to-event outcomes such as mortality is a fundamental task in clinical decision-making, commonly addressed through survival analysis. W