LiWi: Layering in the Wild
arXiv:2605.14552v1 Announce Type: new Abstract: Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains
Knowledge catalogue
arXiv:2605.14552v1 Announce Type: new Abstract: Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains
arXiv:2605.14354v1 Announce Type: new Abstract: We present a new computational framework for detecting and structuring manipulative political narratives. A task that became more important due to the s
arXiv:2605.14202v1 Announce Type: cross Abstract: Malformed, missing, or boundary-value inputs in microservice APIs can cascade across dependent services, threatening reliability. Robustness testing s
arXiv:2605.14186v1 Announce Type: new Abstract: Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succ
arXiv:2603.16659v3 Announce Type: replace Abstract: Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks, including mathematics, code, and structure prediction. The harder b
arXiv:2604.05306v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We st
arXiv:2605.14548v1 Announce Type: new Abstract: Gait recognition, as a promising biometric technology, identifies individuals through their unique walking patterns and offers distinctive advantages in
arXiv:2605.15108v1 Announce Type: cross Abstract: Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging
arXiv:2605.14365v1 Announce Type: cross Abstract: Recent tabular learning benchmarks increasingly show a tight performance cluster rather than a clear hierarchy among leading methods, spanning gradien
arXiv:2508.06202v2 Announce Type: replace-cross Abstract: Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. Howeve
arXiv:2601.21929v2 Announce Type: replace Abstract: Training data attribution (TDA) identifies which training examples most influenced a model's prediction. Influence function methods are a theoretica
arXiv:2505.23912v2 Announce Type: replace-cross Abstract: Hallucination remains a major challenge for the safe and trustworthy deployment of large language models (LLMs) in factual content generation.
arXiv:2605.14874v1 Announce Type: new Abstract: Virtual Try-On (VTON) aims to synthesize photorealistic images of garments precisely aligned with a person's body and pose. Current diffusion-based meth
arXiv:2603.14360v2 Announce Type: replace-cross Abstract: Transformers are highly parallel but are limited to computations in the TC^0 complexity class, excluding tasks such as entity tracking and cod
arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t
arXiv:2602.16898v5 Announce Type: replace-cross Abstract: Task planning for robotic manipulation with large language models (LLMs) is an emerging area. Prior approaches rely on specialized models, fin
arXiv:2605.14606v1 Announce Type: new Abstract: Accurate precipitation nowcasting over extended horizons (0-3 hours) is essential for disaster mitigation and operational decision-making, yet remains a
arXiv:2510.15141v4 Announce Type: replace-cross Abstract: Most existing manifold dimension estimators rely on the assumption that the underlying manifold is locally flat within the neighborhoods under
arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein
arXiv:2601.20173v2 Announce Type: replace Abstract: We present a new nonlinear dimensionality reduction method, MAPLE, that enhances UMAP by improving manifold modeling. MAPLE employs a self-supervise
arXiv:2605.14031v1 Announce Type: cross Abstract: Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositor
arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie
arXiv:2605.14061v1 Announce Type: new Abstract: Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains
arXiv:2605.14304v1 Announce Type: cross Abstract: Compositional generalization in sequential decision-making requires identifying which parts of prior rollouts remain useful for new tasks. Existing me
arXiv:2509.01416v2 Announce Type: replace Abstract: The computational overhead of traditional numerical solvers for partial differential equations (PDEs) remains a critical bottleneck for large-scale
arXiv:2605.14087v1 Announce Type: new Abstract: Large Language Models (LLMs), when trained on web-scale corpora, inherently absorb toxic patterns from their training data. This leads to ``toxic degene
arXiv:2605.14021v1 Announce Type: cross Abstract: Google AI Overviews (AIOs) are arguably the most widely encountered deployment of generative AI, reaching over 2 billion users who may not realize the
arXiv:2605.14744v1 Announce Type: cross Abstract: Large language models in regulated financial workflows are governed by natural-language policies that the same model interprets, creating a principal-
arXiv:2605.13930v1 Announce Type: new Abstract: EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to
arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion
arXiv:2605.14579v1 Announce Type: new Abstract: Accurate medical image segmentation is fundamental to precision medicine, yet robust delineation remains challenging under heterogeneous appearances, am
arXiv:2605.14771v1 Announce Type: new Abstract: MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, plu
arXiv:2511.02271v2 Announce Type: replace Abstract: Medical Report Generation (MRG) is a key part of modern medical diagnostics, as it automatically generates reports from radiological images to reduc
arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r
arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili
arXiv:2605.14421v1 Announce Type: cross Abstract: We introduce MemLineage, a defense for LLM agent memory that attaches both cryptographic provenance and LLM-mediated derivation lineage to every entry
arXiv:2605.15156v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong performance across a wide range of tasks, but remain frozen after pretraining until subsequent updates. Ma
arXiv:2510.15849v2 Announce Type: replace Abstract: Accurate tongue segmentation is crucial for reliable TCM analysis. Supervised models require large annotated datasets, while SAM-family models remai
arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're
arXiv:2605.13919v1 Announce Type: new Abstract: Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods
arXiv:2605.14960v1 Announce Type: cross Abstract: Impossible objects, geometric constructions that humans can perceive but that cannot exist in real life, have been a topic of intrigue in visual arts,
arXiv:2605.14212v1 Announce Type: new Abstract: Automatic multi-agent systems aim to instantiate agent workflows without relying on manually designed or fixed orchestration. However, existing automati
arXiv:2605.15172v1 Announce Type: cross Abstract: Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safe
arXiv:2605.14289v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale capacity by combining specialized experts, but most existing approaches assume centralized access to training da
arXiv:2605.14407v1 Announce Type: new Abstract: The dominant discourse on AI limitations frames the boundary of AI capability as a divide between digital tasks (where AI excels) and physical tasks (wh
arXiv:2605.14966v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have achieved remarkable performance across diverse multimodal tasks, yet they continue to suffer from hallucinat
arXiv:2605.14980v1 Announce Type: cross Abstract: Analyzing microscopy images to extract biological object properties (e.g., their morphological organization, temporal dynamics, and population density
arXiv:2509.14159v3 Announce Type: replace Abstract: As robots become more integrated in society, their ability to coordinate with other robots and humans on multi-modal tasks (those with multiple vali
arXiv:2605.14660v1 Announce Type: new Abstract: Post-Traumatic Stress Disorder (PTSD) is fundamentally a neuroplastic problem traumatic contact events encode over-reactive neural pathways through Hebb
arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s
arXiv:2605.14292v1 Announce Type: cross Abstract: KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavi
arXiv:2605.14568v1 Announce Type: cross Abstract: Context. Behaviour-Driven Development (BDD) software test suites accumulate duplicated step subsequences. Three published refactoring patterns are ava
arXiv:2605.14005v1 Announce Type: new Abstract: Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens a
arXiv:2605.14380v1 Announce Type: new Abstract: Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Au
arXiv:2605.14530v1 Announce Type: new Abstract: Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding fo
arXiv:2605.14664v1 Announce Type: new Abstract: Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the in
arXiv:2605.13849v1 Announce Type: new Abstract: Determining what to eat to satisfy nutritional requirements is one of the oldest optimization problems in operations research, yet existing formulations
arXiv:2509.22746v2 Announce Type: replace Abstract: Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be achieved in particular domains, th
arXiv:2605.15081v1 Announce Type: cross Abstract: The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitiv
arXiv:2510.23477v2 Announce Type: replace Abstract: Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimoda