End-to-end Listen, Look, Speak and Act
arXiv:2510.16756v2 Announce Type: replace-cross Abstract: Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and
Knowledge catalogue
arXiv:2510.16756v2 Announce Type: replace-cross Abstract: Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and
arXiv:2506.02718v2 Announce Type: replace Abstract: Large language models (LLMs) are versatile, yet their deployment in complex real-world settings is limited by static knowledge cutoffs and the diffi
arXiv:2604.18066v1 Announce Type: cross Abstract: Anomaly-based Intrusion Detection Systems (IDSs) ensure protection against malicious attacks on networked systems. While deep learning-based IDSs achi
arXiv:2603.15299v2 Announce Type: replace Abstract: We propose a novel approach which exploits chaos to enhance classification accuracy. Specifically, the available data that need to be classified are
arXiv:2604.18075v1 Announce Type: new Abstract: We investigate recently introduced domain-class incremental learning scenarios for vision-language models (VLMs). Recent works address this challenge us
arXiv:2604.18336v1 Announce Type: cross Abstract: Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth A
arXiv:2412.02904v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized the field of natural language processing with their impressive reasoning and question-answering cap
arXiv:2604.17233v1 Announce Type: new Abstract: Personalized image aesthetics assessment (PIAA) aims to predict an individual user's subjective rating of an image, which requires modeling user-specifi
arXiv:2510.15218v3 Announce Type: replace Abstract: The stacking ensemble combining RF, LightGBM, and DNN performed well on internal test sets, exhibiting an NPV greater than 99.9% even with substanti
arXiv:2501.12119v3 Announce Type: replace-cross Abstract: We introduce ENTIRE, a novel deep learning-based approach for fast and accurate volume rendering time prediction. Predicting rendering time is
arXiv:2510.00861v2 Announce Type: replace Abstract: While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limite
arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable
arXiv:2603.03692v2 Announce Type: replace Abstract: Classifier-Free Guidance (CFG) has established the foundation for guidance mechanisms in diffusion models, showing that well-designed guidance proxi
arXiv:2410.04509v3 Announce Type: replace Abstract: As the field of Multimodal Large Language Models (MLLMs) continues to evolve, their potential to revolutionize artificial intelligence is particular
arXiv:2604.18452v1 Announce Type: cross Abstract: Vision-language modeling is rapidly increasing in popularity with an ever expanding list of available models. In most cases, these vision-language mod
arXiv:2505.15353v3 Announce Type: replace Abstract: Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterog
arXiv:2502.13464v2 Announce Type: replace Abstract: Commonsense plausibility estimation is critical for evaluating language models (LMs), yet existing generative approaches--reliant on likelihoods or
arXiv:2509.11206v4 Announce Type: replace-cross Abstract: Practitioners increasingly rely on Large Language Models (LLMs) to evaluate generative AI outputs through 'LLM-as-a-Judge' approaches. However
arXiv:2604.16744v1 Announce Type: new Abstract: We present a framework for evaluating adaptive personalization of educational reading materials with theory-grounded simulated learners. The system buil
arXiv:2604.16980v1 Announce Type: new Abstract: Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data,
arXiv:2604.16575v1 Announce Type: new Abstract: Unsupervised anomaly detection is widely used to detect Distributed Denial-of-Service (DDoS) attacks in cloud-native 5G networks, yet most studies assum
arXiv:2508.17458v2 Announce Type: replace Abstract: Verbal multiword expressions (VMWEs) remain difficult for machine translation because their meanings are often not recoverable from their component
arXiv:2604.16706v1 Announce Type: cross Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, but this assumption has rarely been validated a
arXiv:2604.18320v1 Announce Type: new Abstract: Self-evolution of multimodal large language models (MLLMs) remains a critical challenge: pseudo-label-based methods suffer from progressive quality degr
arXiv:2601.10306v2 Announce Type: replace-cross Abstract: While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to long-context scenarios is hindered by sparsity of outcome rewards
arXiv:2604.17087v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficie
arXiv:2508.07809v5 Announce Type: replace Abstract: Reinforcement learning with verifiable reward (RLVR) has become a promising paradigm for post-training large language models (LLMs) to improve their
arXiv:2604.17753v1 Announce Type: cross Abstract: Merging multiple Low-Rank Adaptation (LoRA) experts into a single backbone is a promising approach for efficient multi-task deployment. While existing
arXiv:2604.18052v1 Announce Type: cross Abstract: Intrusion detection systems (IDSs) for 5G networks must handle complex, high-volume traffic. Although opaque 'black-box' models can achieve high accur
arXiv:2602.16090v2 Announce Type: replace-cross Abstract: The response of the climate system to increased greenhouse gases and other radiative perturbations is governed by a combination of fast and sl
arXiv:2604.16528v1 Announce Type: new Abstract: Embryo selection is one of multiple crucial steps in in-vitro fertilization, commonly based on morphological assessment by clinical embryologists. Altho
arXiv:2504.04814v3 Announce Type: replace-cross Abstract: Trustworthy artificial intelligence (AI) is essential in healthcare, particularly for high-stakes tasks like medical image segmentation. Expla
arXiv:2512.11108v3 Announce Type: replace Abstract: Good quality explanations strengthen the understanding of language models and data. Feature attribution methods, such as Integrated Gradient, are a
arXiv:2604.17879v1 Announce Type: new Abstract: Camouflaged Object Detection is challenging due to the high degree of similarity between camouflaged objects and their surrounding backgrounds. Current
arXiv:2604.18296v1 Announce Type: new Abstract: Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor
arXiv:2502.13637v2 Announce Type: replace Abstract: Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action with
arXiv:2604.16757v1 Announce Type: new Abstract: The expression of emotions that serve social purposes, such as asserting independence or fostering interdependence, is central to human interactions and
arXiv:2505.22226v2 Announce Type: replace Abstract: Recent theoretical advances reveal that the Hadamard product induces nonlinear representations and implicit high-dimensional mappings for the field
arXiv:2604.18168v1 Announce Type: new Abstract: Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existin
arXiv:2604.16865v1 Announce Type: cross Abstract: In this paper, we consider the problem of extraction of most informative features from time series that are regarded as observed values of stochastic
arXiv:2604.16450v1 Announce Type: cross Abstract: Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess de
arXiv:2604.16610v1 Announce Type: cross Abstract: Machine learning models often inherit biases from historical data, raising critical concerns about fairness and accountability. Conventional fairness
arXiv:2604.16780v1 Announce Type: new Abstract: This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves both representation and predictio
arXiv:2506.12176v5 Announce Type: replace Abstract: In explainable AI, surrogate models are commonly evaluated by their fidelity to a neural network's predictions. Fidelity, however, measures alignmen
arXiv:2601.11886v2 Announce Type: replace Abstract: In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the
arXiv:2512.20182v3 Announce Type: replace Abstract: Recognizing whether outputs from large language models (LLMs) contain faithfulness hallucination is crucial for real-world applications, e.g., retri
arXiv:2604.16522v1 Announce Type: new Abstract: This paper proposes a fast and online method for jointly performing 3D multi-object tracking and pose estimation using multiple monocular cameras. Our a
arXiv:2604.18491v1 Announce Type: new Abstract: Computational Fluid Dynamics (CFD) is central to race-car aerodynamic development, yet its cost -- tens of thousands of core-hours per high-fidelity eva
arXiv:2604.17956v1 Announce Type: new Abstract: Machine learning has become integral to medical research and is increasingly applied in clinical settings to support diagnosis and decision-making; howe
arXiv:2604.16778v1 Announce Type: new Abstract: LLM-powered agents often reason from scratch when presented with a new problem instance and lack automatic mechanisms to transfer learned skills to othe
arXiv:2604.16612v1 Announce Type: new Abstract: Traffic prediction plays a central role in intelligent transportation systems (ITS) by supporting real-time decision-making, congestion management, and
arXiv:2604.16574v1 Announce Type: new Abstract: Federated Learning (FL) faces challenges from client data heterogeneity and resource-constrained mobile devices, which can degrade model accuracy. Perso
arXiv:2510.24942v2 Announce Type: replace-cross Abstract: Despite their impressive performance, vision-language models (VLMs) still struggle on culturally situated inputs. To understand how VLMs proce
arXiv:2511.17171v4 Announce Type: replace Abstract: Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer
arXiv:2604.17919v1 Announce Type: new Abstract: Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However
arXiv:2604.16649v1 Announce Type: new Abstract: Directed energy deposition (DED) produces complex thermo-mechanical responses that can lead to distortion and reduced dimensional accuracy of a manufact
arXiv:2604.17344v1 Announce Type: cross Abstract: When task-specific labels are not available, it becomes difficult to select an embedding model for a specific target corpus. Existing labelless measur
arXiv:2604.17513v1 Announce Type: new Abstract: Simulation frameworks such as Isaac Sim have enabled scalable robot learning for locomotion and rigid-body manipulation; however, contact-rich simulatio
arXiv:2604.17720v1 Announce Type: cross Abstract: Point-based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sa
arXiv:2511.00868v2 Announce Type: replace Abstract: Large Language Model (LLM) serving is increasingly constrained by the growing size of the key-value (KV) cache, which scales with both context lengt