AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
2 Jun 2026

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2606.00684v1 Announce Type: cross Abstract: We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using

LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning

ApplicationsDGX agent

arXiv:2606.01336v1 Announce Type: new Abstract: As real-world applications increasingly require processing inputs of 100k+ tokens, the gap between context length and inference efficiency has become a

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

SafetyDGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

M^3 Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models

ResearchDGX agent

arXiv:2410.12325v2 Announce Type: replace Abstract: In this paper, we study a fundamental design problem in pretraining Large Language Models (LLMs) for low-resource language regimes. Existing works a

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

ResearchDGX agent

arXiv:2606.02004v1 Announce Type: new Abstract: Consumer-price measurement increasingly draws on alternative data sources -- scanner, web-scraped, and transaction/receipt data. A recurring obstacle is

Malaysian English News Decoded: A Linguistic Resource for Named Entity and Relation Extraction

ResearchDGX agent

arXiv:2402.14521v2 Announce Type: replace Abstract: Standard English and Malaysian English exhibit notable differences, posing challenges for natural language processing (NLP) tasks on Malaysian Engli

MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

Model ReleasesDGX agent

arXiv:2505.18614v5 Announce Type: replace Abstract: Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated mu

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

SafetyDGX agent

arXiv:2606.01914v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly atten

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning

Model ReleasesDGX agent

arXiv:2606.01301v1 Announce Type: new Abstract: Hallucinations in medical large language models (LLMs) pose serious risks for clinical decision support, particularly when models must reason over compl

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

Model ReleasesDGX agent

arXiv:2603.20884v2 Announce Type: replace Abstract: To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for pap

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

AgentsDGX agent

arXiv:2602.03318v3 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLM

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

SafetyDGX agent

arXiv:2606.01926v1 Announce Type: new Abstract: Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) app

Model-Based Quality Assessment for Massively Multilingual Parallel Data

Model ReleasesDGX agent

arXiv:2606.00285v1 Announce Type: new Abstract: Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based

Modeling Distinct Human Interaction in Web Agents

AgentsDGX agent

arXiv:2602.17588v3 Announce Type: replace Abstract: Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

Model ReleasesDGX agent

arXiv:2606.00832v1 Announce Type: new Abstract: Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmark

Multi-Agent Computer Use

Model ReleasesDGX agent

arXiv:2606.01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents. This setup is suboptimal for complex long-horizon tasks that benefit

Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony

Model ReleasesDGX agent

arXiv:2512.16401v5 Announce Type: replace Abstract: Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-

NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization

SafetyDGX agent

arXiv:2511.20409v2 Announce Type: replace Abstract: Text normalization methods such as stemming and lemmatization are fundamental components of NLP pipelines. As new normalization tools are developed

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

ResearchDGX agent

arXiv:2606.01148v1 Announce Type: new Abstract: Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support s

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

AgentsDGX agent

arXiv:2606.00820v1 Announce Type: new Abstract: Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that co

Not What, But How: A Communicative Audit of LLM Response Framing

ResearchDGX agent

arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses

OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources

Model ReleasesDGX agent

arXiv:2405.01930v2 Announce Type: replace Abstract: This paper introduces OARelatedWork: a dataset for related work generation from open-access sources. It is the first large-scale multi-document summ

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

ResearchDGX agent

arXiv:2606.00683v1 Announce Type: new Abstract: Recent progress in the development of language models has been defined by scale, with each generation absorbing more of the world's knowledge into its w

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

Model ReleasesDGX agent

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

On the Generalization Gap in Self-Evolving Language Model Reasoning

Model ReleasesDGX agent

arXiv:2606.01075v1 Announce Type: new Abstract: Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself.

On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective

ResearchDGX agent

arXiv:2606.02158v1 Announce Type: new Abstract: AI-generated text increasingly blends with human writing, raising practical risks such as misinformation, academic misuse, and corpora contamination. Wh

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Model ReleasesDGX agent

arXiv:2606.02437v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapt

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

Model ReleasesDGX agent

arXiv:2510.17532v2 Announce Type: replace Abstract: Predicting cancer treatment outcomes requires models that are both accurate and interpretable, particularly in the presence of heterogeneous clinica

PaperVoyager : Building Interactive Web with Visual Language Models

Model ReleasesDGX agent

arXiv:2603.22999v3 Announce Type: replace Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, exist

Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models

Model ReleasesDGX agent

arXiv:2606.00284v1 Announce Type: new Abstract: While continual pretraining~(CPT) is a practical way to extend large language models to new languages, naive finetuning on targeted data erodes existing

Parametric Social Identity Injection and Diversification in Public Opinion Simulation

ApplicationsDGX agent

arXiv:2603.16142v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costl

Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

ResearchDGX agent

arXiv:2606.01469v1 Announce Type: new Abstract: The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually

Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction

ResearchDGX agent

arXiv:2507.18863v2 Announce Type: replace-cross Abstract: Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, s

PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining

ResearchDGX agent

arXiv:2606.01049v1 Announce Type: new Abstract: Large-scale biomedical image-text datasets extracted from scientific literature provide valuable resources for medical multimodal model training. These

PortBERT: Navigating the Depths of Portuguese Language Models

HardwareDGX agent

arXiv:2606.02100v1 Announce Type: new Abstract: Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neg

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya

Model ReleasesDGX agent

arXiv:2604.04937v1 Announce Type: cross Abstract: Large language models produce fluent text but struggle with systematic reasoning, often hallucinating confident but unfounded claims. When Apple resea

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

Local AiDGX agent

arXiv:2606.00523v1 Announce Type: new Abstract: Standard Large Language Models (LLMs) follow a read-then-generate paradigm, causing unnecessary latency and computation. Streaming LLMs alleviate this i

ProtStructQA: A Denotation Threshold in Protein Structural Reasoning

Model ReleasesDGX agent

arXiv:2606.00451v1 Announce Type: new Abstract: Protein-language systems are often evaluated by whether they generate plausible biological text, but a structural question has a sharper semantics: it d

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

Model ReleasesDGX agent

arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode colla

R2-Router: A New Paradigm for LLM Routing with Reasoning

ResearchDGX agent

arXiv:2602.02823v2 Announce Type: replace Abstract: As LLMs proliferate with diverse capabilities and costs, LLM routing has emerged by learning to predict each LLM's quality and cost for a given quer

RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift

TutorialsDGX agent

arXiv:2606.01697v1 Announce Type: new Abstract: Conversational search has become increasingly important in retrieval-augmented generation (RAG) systems, where users interact with AI assistants through

RealityTest: How People Probe AI Identity and Whether Models Disclose It

Model ReleasesDGX agent

arXiv:2606.00168v1 Announce Type: new Abstract: AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mo

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

Model ReleasesDGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

Reconsidering Positional Supervision in Masked Diffusion Language Model Training

SafetyDGX agent

arXiv:2601.22947v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive l

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

SafetyDGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

RenoBench: A Citation Parsing Benchmark

Model ReleasesDGX agent

arXiv:2603.25640v2 Announce Type: replace-cross Abstract: Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, exi

ResMerge: Residual-based Spectral Merging of Large Language Models

ResearchDGX agent

arXiv:2606.02252v1 Announce Type: new Abstract: Model merging offers a training-free way to combine multiple post-trained expert models, but merging experts obtained through reinforcement learning (RL

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

Model ReleasesDGX agent

arXiv:2606.01923v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit 'contextual disregard' when faced with input evidence that conflicts with their internal parametric memo

Revise, Don't Freeze: Sampler-Matched Training for Self-Correcting Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2606.01026v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving th

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.01600v1 Announce Type: cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instruc

Robust Asynchronous Planning via Auto-Formalization

ApplicationsDGX agent

arXiv:2606.00981v1 Announce Type: new Abstract: LLMs can plan by either generating action sequences directly as a Planner or translating tasks into domain specific language for an external solver as a

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation

TutorialsDGX agent

arXiv:2606.00628v1 Announce Type: new Abstract: Self-distillation improves learning efficiency by rewriting reference answers as training data that better matches the model's own distribution. However

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors

ResearchDGX agent

arXiv:2606.00460v1 Announce Type: new Abstract: Speech-aware large language models often generalize poorly to out-of-domain settings. We propose SALSA (Speech-Aware LLM Adaptation via Learned Steering

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

Model ReleasesDGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

Model ReleasesDGX agent

arXiv:2606.00579v1 Announce Type: new Abstract: As multimodal LLMs increasingly target video and audio, it is often assumed that such tasks require native omnimodal models. We show that this is not al

SARA: Stress Test Reasoning in Audio Deepfake Detection

ResearchDGX agent

arXiv:2601.03615v2 Announce Type: replace Abstract: Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond extit{black-box} classifiers

Scaling Agentic Capabilities via Grounded Interaction Synthesis

AgentsDGX agent

arXiv:2606.02001v1 Announce Type: new Abstract: General agentic intelligence hinges on the ability to interact with diverse real-world tools to complete complex tasks, a capability fundamentally tied

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

AgentsDGX agent

arXiv:2602.12984v2 Announce Type: replace Abstract: Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely ov

Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs

ResearchDGX agent

arXiv:2510.08825v2 Announce Type: replace Abstract: Large language models (LLMs) augmented with knowledge graphs (KGs) offer a promising approach for knowledge-intensive reasoning. Central to this app

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

ResearchDGX agent

arXiv:2510.24870v2 Announce Type: replace Abstract: We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a preval

← Previous
1…4950515253…129
Next →