Teaching an Agent to Sketch One Part at a Time
arXiv:2603.19500v2 Announce Type: replace Abstract: We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel mu
Knowledge catalogue
arXiv:2603.19500v2 Announce Type: replace Abstract: We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel mu
arXiv:2604.22237v1 Announce Type: cross Abstract: Diagnosing student problem behaviors requires teachers to synthesize multifaceted information, identify behavioral categories, and plan intervention s
arXiv:2510.07632v2 Announce Type: replace Abstract: Frontier AI models have achieved remarkable progress, yet recent studies suggest they struggle with compositional reasoning, often performing at or
arXiv:2604.21938v1 Announce Type: cross Abstract: Embodied AI is widely discussed as a job-displacement problem. The deeper risk, however, is governance lag: the inability of public institutions to ke
arXiv:2505.20435v3 Announce Type: replace-cross Abstract: Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlook
arXiv:2506.17299v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their v
arXiv:2604.22260v1 Announce Type: cross Abstract: Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While rece
arXiv:2512.20761v3 Announce Type: replace-cross Abstract: Time Series Foundation Models (TSFMs) are transforming the field of forecasting. However, evaluating them on historical data is increasingly d
arXiv:2604.22209v1 Announce Type: cross Abstract: Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each
arXiv:2604.21999v1 Announce Type: cross Abstract: We study learned memory tokens as computational scratchpad for a single-block Universal Transformer (UT) with Adaptive Computation Time (ACT) on Sudok
arXiv:2508.06165v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for know
arXiv:2604.22215v1 Announce Type: cross Abstract: Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models
arXiv:2603.16663v5 Announce Type: replace-cross Abstract: The AIED community envisions AI evolving 'from tools to teammates,' yet most research still examines AI agents primarily through one-on-one hu
arXiv:2604.22153v1 Announce Type: cross Abstract: When you ask an AI assistant for advice about your career, your marriage, or a conflict with your family, does it give you the same answer regardless
arXiv:2604.22273v1 Announce Type: new Abstract: Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear. We frame self-correcti
arXiv:2510.21285v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harm
arXiv:2604.22102v1 Announce Type: cross Abstract: Many robotic tasks are unforgiving; a single mistake in a dynamic throw can lead to unacceptable delays or unrecoverable failure. To mitigate this, we
arXiv:2509.03294v3 Announce Type: replace-cross Abstract: The increasing availability of personal data has enabled significant advances in fields such as machine learning, healthcare, and cybersecurit
arXiv:2604.21028v1 Announce Type: cross Abstract: The increasing frequency and severity of global flood events highlights the need for the development of rapid and reliable flood prediction tools. Thi
arXiv:2604.21579v1 Announce Type: cross Abstract: LLM-based automated program repair (APR) techniques have shown promising results in reducing debugging costs. However, prior results can be affected b
arXiv:2604.21891v1 Announce Type: cross Abstract: Maintaining instantaneous balance between electricity supply and demand is critical for reliability and grid instability. System operators achieve thi
arXiv:2604.21885v1 Announce Type: cross Abstract: Event extraction is essential for event understanding and analysis. It supports tasks such as document summarization and decision-making in emergency
arXiv:2510.08814v2 Announce Type: replace-cross Abstract: We present a proof architecture for (P neq NP) based on an upper--lower clash in polytime-capped conditional description length. We construct
arXiv:2604.21903v1 Announce Type: cross Abstract: Deep-learning video super-resolution has progressed rapidly, but climate applications typically super-resolve (increase resolution) either space or ti
arXiv:2604.21030v1 Announce Type: cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decision-making
arXiv:2604.20915v1 Announce Type: cross Abstract: Transformers suffer from a high computational cost that grows with sequence length for self-attention, making inference in long streams prohibited by
arXiv:2510.18091v2 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) partition input images into uniformly sized patches regardless of their content, resulting in long input sequence l
arXiv:2604.21044v1 Announce Type: new Abstract: In some complex domains, certain problem-specific decompositions can provide advantages over monolithic designs by enabling comprehension and specificat
arXiv:2604.20932v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems are increasingly deployed in sensitive domains such as healthcare and law, where they rely on private, do
arXiv:2604.21159v1 Announce Type: cross Abstract: Many approaches to LLM red-teaming leverage an attacker LLM to discover jailbreaks against a target. Several of them task the attacker with identifyin
arXiv:2604.21018v1 Announce Type: new Abstract: While scaling test-time compute can substantially improve model performance, existing approaches either rely on static compute allocation or sample from
arXiv:2511.04638v5 Announce Type: replace-cross Abstract: A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to under
arXiv:2604.21879v1 Announce Type: cross Abstract: The ability of generative AI (GenAI) methods to photorealistically alter camera images has raised awareness about the authenticity of images shared on
arXiv:2604.20846v1 Announce Type: cross Abstract: Next point-of-interest (POI) recommendation requires modeling user mobility as a spatiotemporal sequence, where different behavioral factors may evolv
arXiv:2604.21310v1 Announce Type: cross Abstract: Deep learning has emerged as a powerful approach for malware detection, demonstrating impressive accuracy across various data representations. However
arXiv:2604.21725v1 Announce Type: cross Abstract: LLM agents increasingly operate in open-ended environments spanning hundreds of sequential episodes, yet they remain largely stateless: each task is s
arXiv:2601.18491v2 Announce Type: replace Abstract: The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current gua
arXiv:2604.21744v1 Announce Type: cross Abstract: The capabilities of AI-assisted coding are progressing at breakneck speed. Chat-based vibe coding has evolved into fully fledged AI-assisted, agentic
arXiv:2604.21154v1 Announce Type: new Abstract: At-home physiotherapy compliance remains critically low due to a lack of personalized supervision and dynamic feedback. Existing digital health solution
arXiv:2604.21129v1 Announce Type: cross Abstract: Current blockchain Layer 2 solutions, including Optimism, Arbitrum, zkSync, and their derivatives, optimize for human-initiated financial transactions
arXiv:2511.23159v2 Announce Type: replace-cross Abstract: Vibe coding, the much-touted use of AI techniques for programming, faces two overwhelming obstacles: the difficulty of specifying goals ('prom
arXiv:2604.21103v1 Announce Type: new Abstract: Governments are increasingly interested in using AI to make administrative decisions cheaper, more scalable, and more consistent. But for probabilistic
arXiv:2604.21446v1 Announce Type: new Abstract: We present AI-Gram, a live platform enabling image-based interactions, to study social dynamics in a fully autonomous multi-agent visual network where a
arXiv:2507.15753v2 Announce Type: replace-cross Abstract: Generative machine learning models have revolutionized material discovery by capturing complex structure-property relationships, yet extending
arXiv:2604.21209v1 Announce Type: new Abstract: Online reviews have played a pivotal role in consumers' decision-making processes. Existing research has highlighted the significant impact of manageria
arXiv:2604.21827v1 Announce Type: new Abstract: Modern AI assistants are trained to follow instructions, implicitly assuming that users can clearly articulate their goals and the kind of assistance th
arXiv:2602.13211v2 Announce Type: replace-cross Abstract: Compared with IP multicast, Overlay Multicast (OM) offers better compatibility and flexible deployment in heterogeneous, cross-domain networks
arXiv:2502.04416v3 Announce Type: replace-cross Abstract: Scaling large language models (LLMs) improves performance but significantly increases inference costs, with feed-forward networks (FFNs) consu
arXiv:2604.20862v1 Announce Type: new Abstract: The automation system for Course of Action (CoA) planning is an essential element in future warfare. As maneuver speeds increase, surveillance ranges ex
arXiv:2604.21529v1 Announce Type: cross Abstract: Applying the concept of controlled self-organization in agent-based Cyber-Physical Energy Systems (CPES) is a promising approach to ensure system robu
arXiv:2604.20850v1 Announce Type: cross Abstract: Dense retrieval systems rank passages by embedding similarity to a query, but multi-hop questions require passages that are associatively related thro
arXiv:2603.01170v2 Announce Type: replace-cross Abstract: This work presents ATLAS, an LLM-driven framework that bridges standardized threat modeling and property-based formal verification for System-
arXiv:2604.20844v1 Announce Type: cross Abstract: Recent GraphRAG methods integrate graph structures into text indexing and retrieval, using knowledge graph triples to connect text chunks, thereby imp
arXiv:2604.21530v1 Announce Type: cross Abstract: Lung adenocarcinoma (LUAD) grading depends on accurately identifying growth patterns, which are indicators of prognosis and can influence treatment de
arXiv:2305.01626v4 Announce Type: replace-cross Abstract: Computational models of syntax are predominantly text-based. Here we propose that the most basic first step in the evolution of syntax can be
arXiv:2604.21083v1 Announce Type: cross Abstract: Third-party Large Language Model (LLM) API gateways are rapidly emerging as unified access points to models offered by multiple vendors. However, the
arXiv:2604.21344v1 Announce Type: cross Abstract: Charts are widely used to present complex information. Deriving meaningful insights in real-world contexts often requires interpreting multiple relate
arXiv:2604.20906v1 Announce Type: cross Abstract: The rapid growth of scientific software has created practical barriers for bioinformatics research. Although powerful statistical, artificial intellig
arXiv:2604.21508v1 Announce Type: new Abstract: Protein-ligand bioactivity data published in the literature are essential for drug discovery, yet manual curation struggles to keep pace with rapidly gr
arXiv:2604.21854v1 Announce Type: new Abstract: Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Go