AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
29 Apr 2026

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering

Model ReleasesDGX agent

arXiv:2601.12248v2 Announce Type: replace-cross Abstract: Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks m

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation

ResearchDGX agent

arXiv:2604.25702v1 Announce Type: new Abstract: Contemporary neural machine translation (NMT) systems are almost exclusively built by training on supervised parallel data. Despite the tremendous progr

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

Barriers to Universal Reasoning With Transformers (And How to Overcome Them)

TutorialsDGX agent

arXiv:2604.25800v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing comple

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

Model ReleasesDGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

Model ReleasesDGX agent

arXiv:2604.24955v1 Announce Type: new Abstract: As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

Model ReleasesDGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

Benchmarking Logistic Regression, SVM, and LightGBM Against BiLSTM with Attention for Sentiment Analysis on Indonesian Product Reviews

ResearchDGX agent

arXiv:2604.25452v1 Announce Type: new Abstract: Sentiment analysis of product reviews on e-commerce platforms plays a critical role in automatically understanding customer satisfaction and providing a

Benchmarking PyCaret AutoML Against IndoBERT Fine-Tuning for Sentiment Analysis on Indonesian IKN Twitter Data

ResearchDGX agent

arXiv:2604.25392v1 Announce Type: new Abstract: This paper benchmarks a classical machine learning approach based on PyCaret AutoML against a deep learning approach based on IndoBERT fine-tuning for b

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Model ReleasesDGX agent

arXiv:2509.09708v3 Announce Type: replace Abstract: Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour re

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Model ReleasesDGX agent

arXiv:2512.12087v3 Announce Type: replace Abstract: The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation

ResearchDGX agent

arXiv:2604.25580v1 Announce Type: new Abstract: The closure of Perspective API at the end of 2026 discards what has functioned as the de facto standard for automated toxicity measurement in NLP, CSS,

CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation

Model ReleasesDGX agent

arXiv:2604.25774v1 Announce Type: new Abstract: Accurate nutrient estimation from unstructured recipe text is an important yet challenging problem in dietary monitoring, due to ambiguous ingredient te

Cheaper, Better, Faster, Stronger: Robust Text-to-SQL without Chain-of-Thought or Fine-Tuning

Model ReleasesDGX agent

arXiv:2505.14174v2 Announce Type: replace Abstract: LLMs are effective at code generation tasks like text-to-SQL, but is it worth the cost? Many state-of-the-art approaches use non-task-specific LLM t

Citation Failure: Definition, Analysis and Efficient Mitigation

Model ReleasesDGX agent

arXiv:2510.20303v3 Announce Type: replace Abstract: Citations from LLM-based RAG systems are supposed to simplify response verification. However, this goal is undermined in cases of citation failure,

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding

ResearchDGX agent

arXiv:2602.01785v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational eff

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest

AgentsDGX agent

arXiv:2604.25088v1 Announce Type: cross Abstract: Language Model (LM)-based agents remain largely untested in mixed-motive settings where agents must leverage short-term cooperation for long-term comp

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

SafetyDGX agent

arXiv:2604.25676v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval space, typically via query or document translation or m

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

Model ReleasesDGX agent

arXiv:2603.25268v2 Announce Type: replace Abstract: We introduce CRAFT, a multi-agent benchmark for evaluating pragmatic communication in large language models under strict partial information. In thi

CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.25182v1 Announce Type: new Abstract: A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Aug

Cross-Lingual Jailbreak Detection via Semantic Codebooks

Model ReleasesDGX agent

arXiv:2604.25716v1 Announce Type: new Abstract: Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment.

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

Model ReleasesDGX agent

arXiv:2604.25318v1 Announce Type: cross Abstract: Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative

Diagnosis, Bad Planning & Reasoning. Treatment, SCOPE -- Planning for Hybrid Querying over Clinical Trial Data

AgentsDGX agent

arXiv:2604.25120v1 Announce Type: new Abstract: We study clinical trial table reasoning, where answers are not directly stored in visible cells but must be reasoned from semantic understanding through

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference

ResearchDGX agent

arXiv:2510.19669v4 Announce Type: replace Abstract: Recent reasoning Large Language Models (LLMs) demonstrate remarkable problem-solving abilities but often generate long thinking traces whose utility

Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives

ResearchDGX agent

arXiv:2604.25423v1 Announce Type: new Abstract: Do large language models (LLMs) truly acquire embodied cognition and cultural conventions from text? We introduce demonstratives, fundamental spatial ex

Doing More With Less: Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.25098v1 Announce Type: cross Abstract: While current Large Language Models (LLMs) exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), their massive parameter

Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination

Model ReleasesDGX agent

arXiv:2604.24978v1 Announce Type: new Abstract: Enterprise deep research often fails to produce decision-ready reports due to uneven information coverage, context explosion, and premature stopping. We

DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams

Model ReleasesDGX agent

arXiv:2604.25231v1 Announce Type: cross Abstract: Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics

Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

Local AiDGX agent

arXiv:2604.25039v1 Announce Type: new Abstract: Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still strugg

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

Model ReleasesDGX agent

arXiv:2604.25914v1 Announce Type: new Abstract: Real-world data visualization (DV) requires native environmental grounding, cross-platform evolution, and proactive intent alignment. Yet, existing benc

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

Local AiDGX agent

arXiv:2604.24972v1 Announce Type: new Abstract: Clinical abnormality grounding for rare diseases is often hindered by data scarcity, making supervised fine-tuning impractical and single-pass inference

Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR

ResearchDGX agent

arXiv:2604.24770v1 Announce Type: new Abstract: Despite recent progress in automatic speech recognition (ASR), elderly ASR (EASR) remains challenging due to limited training data and the distinct acou

Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis

Model ReleasesDGX agent

arXiv:2603.16877v2 Announce Type: replace Abstract: Financial analysts face significant challenges extracting information from lengthy 10-K reports, which often exceed 100 pages. This paper presents a

Exploring Reasoning Reward Model for Agents

Model ReleasesDGX agent

arXiv:2601.22154v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use. However

Faithful Autoformalization via Roundtrip Verification and Repair

Model ReleasesDGX agent

arXiv:2604.25031v1 Announce Type: new Abstract: When an LLM formalizes natural language, how do we know the output is faithful? We propose a roundtrip verification approach which does not require grou

Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models

Model ReleasesDGX agent

arXiv:2604.25313v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) models frequently produce answers grounded in parametric memory rather than the retrieved context, undermining the

FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments

Model ReleasesDGX agent

arXiv:2604.25135v1 Announce Type: new Abstract: Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environme

Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment

SafetyDGX agent

arXiv:2604.25136v1 Announce Type: new Abstract: We propose Frictive Policy Optimization (FPO), a framework for learning language model policies that regulate not only what to say, but when and how to

From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems

ResearchDGX agent

arXiv:2507.07847v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a crucial framework in natural language processing (NLP), improving factual consistency and redu

From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support

ResearchDGX agent

arXiv:2604.25525v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used not only for instrumental tasks, but as always-available and non-judgmental confidants for emotional

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

Model ReleasesDGX agent

arXiv:2510.18030v2 Announce Type: replace Abstract: Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectu

From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs

ResearchDGX agent

arXiv:2604.25866v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in emotionally sensitive human-AI applications, yet little is known about how emotion recognition is

From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation

Local AiDGX agent

arXiv:2604.25482v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for narrative generation, but their use in complex, multi-layered role-playing game (RPG) world

G-Loss: Graph-Guided Fine-Tuning of Language Models

Model ReleasesDGX agent

arXiv:2604.25853v1 Announce Type: new Abstract: Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

Model ReleasesDGX agent

arXiv:2604.24929v1 Announce Type: new Abstract: Agent benchmarks remain largely English-centric, while their multilingual versions are often built with machine translation (MT) and limited post-editin

Generative AI Carries Non-Democratic Biases and Stereotypes: Representation of Women, Black Individuals, Age Groups, and People with Disability in AI-Generated Images across Occupations

SafetyDGX agent

arXiv:2409.13869v2 Announce Type: replace-cross Abstract: In this study, I investigate how generative artificial intelligence (AI) systems reproduce and reinforce societal biases, with a specific focu

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning

SafetyDGX agent

arXiv:2603.01070v2 Announce Type: replace Abstract: Solving complex geometric problems inherently requires interleaved reasoning: a tight alternation between constructing diagrams and performing logic

Images Amplify Misinformation Sharing in Vision-Language Models

Model ReleasesDGX agent

arXiv:2505.13302v2 Announce Type: replace Abstract: As language and vision-language models (VLMs) become central to information access and online interaction, concerns grow about their potential to am

Improving LLM Predictions via Inter-Layer Structural Encoders

Model ReleasesDGX agent

arXiv:2603.22665v2 Announce Type: replace Abstract: The standard practice in Large Language Models (LLMs) is to base predictions on final-layer representations. However, intermediate layers encode com

Independent-Component-Based Encoding Models of Brain Activity During Story Comprehension

ResearchDGX agent

arXiv:2604.24942v1 Announce Type: new Abstract: Encoding models provide a powerful framework for linking continuous stimulus features to neural activity; however, traditional voxelwise approaches are

Intrinsic Mutual Information as a Modulator for Preference Optimization

ResearchDGX agent

arXiv:2604.24804v1 Announce Type: cross Abstract: Offline preference optimization methods, such as Direct Preference Optimization (DPO), offer significant advantages in aligning Large Language Models

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?

ResearchDGX agent

arXiv:2507.15707v2 Announce Type: replace Abstract: Large Language Models (LLMs) have been evaluated using diverse question types, e.g., multiple-choice, true/false, and short/long answers. This study

Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility

Model ReleasesDGX agent

arXiv:2507.12553v3 Announce Type: replace Abstract: Language models (LMs) are used for a diverse range of tasks, from question answering to writing fantastical stories. In order to reliably accomplish

jina-embeddings-v5-text: Task-Targeted Embedding Distillation

Model ReleasesDGX agent

arXiv:2602.15547v2 Announce Type: replace Abstract: Text embedding models are widely used for semantic similarity tasks, including information retrieval, clustering, and classification. General-purpos

Korean aegyo speech shows systematic F1 increase to signal childlike qualities

ResearchDGX agent

arXiv:2604.25133v1 Announce Type: new Abstract: Korean aegyo is a socially recognized childlike speaking style used predominantly in romantic interactions among adults. This study examined vowel space

Language corpora for the Dutch medical domain

ResearchDGX agent

arXiv:2604.25374v1 Announce Type: new Abstract: extbf{Background:} Dutch medical corpora are scarce, limiting NLP development. extbf{Methods:} We translated English datasets, identified medical text i

Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators

ResearchDGX agent

arXiv:2503.06778v3 Announce Type: replace Abstract: Event annotation is important for identifying market changes, monitoring breaking news, and understanding sociological trends. Although expert annot

Large Language Models Explore by Latent Distilling

Model ReleasesDGX agent

arXiv:2604.24927v1 Announce Type: new Abstract: Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yields surface-le

Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models

SafetyDGX agent

arXiv:2512.20677v4 Announce Type: replace-cross Abstract: The increasing deployment of large language models (LLMs) in safety-critical applications raises fundamental challenges in systematically eval

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

SafetyDGX agent

arXiv:2604.25296v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional

← Previous
1…9495969798…129
Next →