Agentic Insight Generation in VSM Simulations
arXiv:2604.12421v1 Announce Type: new Abstract: Extracting actionable insights from complex value stream map simulations can be challenging, time-consuming, and error-prone. Recent advances in large l
Knowledge catalogue
arXiv:2604.12421v1 Announce Type: new Abstract: Extracting actionable insights from complex value stream map simulations can be challenging, time-consuming, and error-prone. Recent advances in large l
arXiv:2604.12179v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have improved their ability to process extended conversational contexts, yet fine-tuning and evaluat
arXiv:2604.06812v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated impressive capabilities in long-form generation, yet their application is hindered by the hallucinati
arXiv:2604.12162v1 Announce Type: new Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Exi
arXiv:2602.18899v3 Announce Type: replace-cross Abstract: Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplor
arXiv:2604.12196v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate multiple candidate responses for a given prompt, yet selecting the most reliable one remains challengin
arXiv:2604.12471v1 Announce Type: cross Abstract: Scientific novelty drives advances at the research frontier, yet it is also associated with heightened uncertainty and potential resistance from incum
arXiv:2604.12506v1 Announce Type: new Abstract: Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently u
arXiv:2604.12491v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed for tabular question answering, yet calibration on structured data is largely unstudied. This pap
arXiv:2602.05971v2 Announce Type: replace Abstract: Semantic representations can be framed as a structured, dynamic knowledge space through which humans navigate to retrieve and manipulate meaning. To
arXiv:2604.05821v2 Announce Type: replace Abstract: Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less conside
arXiv:2604.12268v1 Announce Type: cross Abstract: Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear.
arXiv:2601.11047v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities but often grapple with reliability challenges like hallucinations.
arXiv:2604.12359v1 Announce Type: cross Abstract: Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain atta
arXiv:2604.12312v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex
arXiv:2604.12308v1 Announce Type: new Abstract: Individuals' concerns about data privacy and AI safety are highly contextualized and extend beyond sensitive patterns. Addressing these issues requires
arXiv:2510.19644v2 Announce Type: replace Abstract: Retrieval-augmented generation has emerged as one of the most effective approaches for code completion enhancement, especially when repository-level
arXiv:2604.12426v1 Announce Type: cross Abstract: We investigate whether transformers use their depth adaptively across tasks of increasing difficulty. Using a controlled multi-hop relational reasonin
arXiv:2604.12659v1 Announce Type: cross Abstract: Vision-language models(VLMs) are increasingly applied to visual stock price forecasting, yet existing benchmarks inadequately evaluate their understan
arXiv:2409.06679v3 Announce Type: replace Abstract: Processing long contexts is increasingly important for Large Language Models (LLMs) in tasks like multi-turn dialogues, code generation, and documen
arXiv:2604.05546v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) enable sophisticated reasoning over images and videos, yet their inference is hindered by a systemic efficiency
arXiv:2604.12047v1 Announce Type: new Abstract: PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, table
arXiv:2604.12518v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) integrates heterogeneous text, audio, and visual signals to infer human emotions. While recent approaches leverage c
arXiv:2510.03323v2 Announce Type: replace Abstract: Integrating textual graphs into Large Language Models (LLMs) is promising for complex graph-based QA. However, a key bottleneck is retrieving inform
arXiv:2510.09536v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -- naturally introducing typographi
arXiv:2604.12776v1 Announce Type: new Abstract: Realizing endogenous narrative evolution in LLM-based multi-agent systems is hindered by the inherent stochasticity of generative emergence. In particul
arXiv:2604.12559v1 Announce Type: new Abstract: Unstructured model editing aims to update models with real-world text, yet existing methods often memorize text holistically without reliable fine-grain
arXiv:2604.12666v1 Announce Type: cross Abstract: Text-based web agents offer computational efficiency for autonomous web navigation, yet developing robust agents remains challenging due to the noisy
arXiv:2604.12385v1 Announce Type: new Abstract: Multi-turn dialogue is the predominant form of interaction with large language models (LLMs). While LLM routing is effective in single-turn settings, ex
arXiv:2604.12748v1 Announce Type: new Abstract: Although large language models (LLMs) excel in complex reasoning tasks, they suffer from severe causal hallucination in event causality identification (
arXiv:2604.12630v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Rec
arXiv:2410.23728v3 Announce Type: replace Abstract: With the increasing quality and spread of LLM assistants, the amount of generated content is growing rapidly. In many cases and tasks, such texts ar
arXiv:2604.12442v1 Announce Type: new Abstract: In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typi
arXiv:2604.12978v1 Announce Type: new Abstract: Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cl
arXiv:2604.02830v2 Announce Type: replace Abstract: Detecting whether a model's internal knowledge is sufficient to correctly answer a given question is a fundamental challenge in deploying responsibl
arXiv:2506.01256v4 Announce Type: replace-cross Abstract: Forced alignment is a common tool to align audio with orthographic and phonetic transcriptions. Most forced alignment tools provide only point
arXiv:2604.05643v2 Announce Type: replace Abstract: Extending CoT through RL has been widely used to enhance the reasoning capabilities of LLMs. However, due to the sparsity of reward signals, it can
arXiv:2604.12843v1 Announce Type: new Abstract: The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are
arXiv:2603.20640v2 Announce Type: replace Abstract: Multi-Agent Debate has emerged as a promising framework for improving the reasoning quality of large language models through iterative inter-agent c
arXiv:2604.12721v1 Announce Type: new Abstract: Clinical case formulation organizes patient symptoms and psychosocial factors into causal models, often using the 5P framework. However, constructing su
arXiv:2603.06552v2 Announce Type: replace Abstract: This paper describes the KCLarity team's participation in CLARITY, a shared task at SemEval 2026 on classifying ambiguity and evasion techniques in
arXiv:2604.12185v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models by grounding outputs in retrieved knowledge. However, existing RAG methods including
arXiv:2604.12397v1 Announce Type: new Abstract: Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that hu
arXiv:2604.12452v1 Announce Type: new Abstract: Large language models (LLMs) face significant challenges in processing long contexts due to the linear growth of the key-value (KV) cache and quadratic
arXiv:2601.14004v4 Announce Type: replace Abstract: Mechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However,
arXiv:2604.12056v1 Announce Type: new Abstract: Block-wise diffusion language models (DLMs) generate multiple tokens in any order, offering a promising alternative to the autoregressive decoding pipel
arXiv:2604.12373v1 Announce Type: new Abstract: Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether larg
arXiv:2604.12479v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major cha
arXiv:2604.12919v1 Announce Type: new Abstract: Metonymy and metaphor often co-occur in natural language, yet computational work has studied them largely in isolation. We introduce a framework that tr
arXiv:2604.12928v1 Announce Type: new Abstract: Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguis
arXiv:2604.12633v1 Announce Type: new Abstract: Emotion classification in multilingual settings remains constrained by the scarcity of annotated data: existing corpora are predominantly English, singl
arXiv:2604.12766v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) typically relies on a flat retrieval paradigm that maps queries directly to static, isolated text segments. This ap
arXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.
arXiv:2512.13961v2 Announce Type: replace Abstract: We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets
arXiv:2604.00136v2 Announce Type: replace-cross Abstract: Multi-model LLM serving operates in a non-stationary, noisy environment: providers revise pricing, model quality can shift or regress without
arXiv:2507.06448v5 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endowing Large Language Models (LLMs) with ro
arXiv:2601.19917v2 Announce Type: replace Abstract: Strategic planning is critical for multi-step reasoning, yet compact Large Language Models (LLMs) often lack the capacity to formulate global strate
arXiv:2604.12995v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability t
arXiv:2604.12378v1 Announce Type: new Abstract: Despite advances in multilingual capabilities, most large language models (LLMs) remain English-centric in their training and, crucially, in their produ
arXiv:2604.12195v1 Announce Type: new Abstract: Work in cognitive science and artificial intelligence has suggested that exposing learning agents to traces of interaction between multiple individuals