AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
16 Apr 2026

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

Model ReleasesDGX agent

arXiv:2601.15550v3 Announce Type: replace Abstract: Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commo

Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic

Model ReleasesDGX agent

arXiv:2604.13065v1 Announce Type: new Abstract: LLMs can execute every step of chain-of-thought reasoning correctly and still produce wrong final answers. We introduce the Novel Operator Test, a bench

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.14121v1 Announce Type: new Abstract: LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, unde

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

Model ReleasesDGX agent

arXiv:2405.19088v3 Announce Type: replace Abstract: Recent advancements in large multimodal language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models sti

Curation of a Palaeohispanic Dataset for Machine Learning

ResearchDGX agent

arXiv:2604.13070v1 Announce Type: new Abstract: Palaeohispanic languages are those spoken in the Iberian Peninsula before the arrival of the Romans in the 3rd Century B.C. Their study was really put o

Debate to Align: Reliable Entity Alignment through Two-Stage Multi-Agent Debate

SafetyDGX agent

arXiv:2604.13551v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify entities referring to the same real-world object across different knowledge graphs (KGs). Recent approaches based

Deep Learning Based Amharic Chatbot for FAQs in Universities

ResearchDGX agent

arXiv:2402.01720v3 Announce Type: replace-cross Abstract: University students often spend a considerable amount of time seeking answers to common questions from administrators or teachers. This can be

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

Model ReleasesDGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

Model ReleasesDGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

Diffusion Language Models for Speech Recognition

TutorialsDGX agent

arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a

Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

Model ReleasesDGX agent

arXiv:2604.13899v1 Announce Type: new Abstract: Instruction-tuned LLMs can annotate thousands of instances from a short prompt at negligible cost. This raises two questions for active learning (AL): c

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

SafetyDGX agent

arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existin

Document-tuning for robust alignment to animals

Model ReleasesDGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

Dual-Enhancement Product Bundling: Bridging Interactive Graph and Large Language Model

ResearchDGX agent

arXiv:2604.14030v1 Announce Type: new Abstract: Product bundling boosts e-commerce revenue by recommending complementary item combinations. However, existing methods face two critical challenges: (1)

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints

ResearchDGX agent

arXiv:2604.13371v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly described as possessing strong reasoning capabilities, supported by high performance on mathematical, logi

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

ResearchDGX agent

arXiv:2604.13286v1 Announce Type: new Abstract: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to p

Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen

Model ReleasesDGX agent

arXiv:2602.24119v2 Announce Type: replace Abstract: Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and be

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

Model ReleasesDGX agent

arXiv:2604.13232v1 Announce Type: new Abstract: This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

Model ReleasesDGX agent

arXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste

EVE: A Domain-Specific LLM Framework for Earth Intelligence

Model ReleasesDGX agent

arXiv:2604.13071v1 Announce Type: new Abstract: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Int

Exposia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Model ReleasesDGX agent

arXiv:2601.06536v2 Announce Type: replace Abstract: We present Exposia, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded com

ExpSeek: Self-Triggered Experience Seeking for Web Agents

Model ReleasesDGX agent

arXiv:2601.08605v2 Announce Type: replace Abstract: Experience intervention in web agents emerges as a promising technical paradigm, enhancing agent interaction capabilities by providing valuable insi

F-Actor: Controllable Conversational Behaviour in Full-Duplex Models

Model ReleasesDGX agent

arXiv:2601.11329v3 Announce Type: replace Abstract: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

Model ReleasesDGX agent

arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi

Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

Local AiDGX agent

arXiv:2508.08791v3 Announce Type: replace Abstract: Effective tool use is essential for large language models (LLMs) to interact with their environment. However, progress is limited by the lack of eff

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

Model ReleasesDGX agent

arXiv:2511.21760v3 Announce Type: replace Abstract: Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capabi

Foresight Optimization for Strategic Reasoning in Large Language Models

SafetyDGX agent

arXiv:2604.13592v1 Announce Type: new Abstract: Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-ba

Form Without Function: Agent Social Behavior in the Moltbook Network

AgentsDGX agent

arXiv:2604.13052v1 Announce Type: cross Abstract: Moltbook is a social network where every participant is an AI agent. We analyze 1,312,238 posts, 6.7~million comments, and over 120,000 agent profiles

From Anchors to Supervision: Memory-Graph Guided Corpus-Free Unlearning for Large Language Models

Local AiDGX agent

arXiv:2604.13777v1 Announce Type: new Abstract: Large language models (LLMs) may memorize sensitive or copyrighted content, raising significant privacy and legal concerns. While machine unlearning has

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Model ReleasesDGX agent

arXiv:2604.14137v1 Announce Type: new Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'':

From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning

ResearchDGX agent

arXiv:2604.13398v1 Announce Type: new Abstract: While Aspect-based Sentiment Analysis (ABSA) systems have achieved high accuracy in identifying sentiment polarities, they often operate as 'black boxes

From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space

SafetyDGX agent

arXiv:2604.14142v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), it

From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines

ApplicationsDGX agent

arXiv:2604.13468v1 Announce Type: cross Abstract: Generative information retrieval (GenIR) formulates the retrieval process as a text-to-text generation task, leveraging the vast knowledge of large la

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

SafetyDGX agent

arXiv:2604.13067v1 Announce Type: cross Abstract: SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations ofte

From Weights to Activations: Is Steering the Next Frontier of Adaptation?

Model ReleasesDGX agent

arXiv:2604.14090v1 Announce Type: new Abstract: Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficie

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

SafetyDGX agent

arXiv:2604.14053v1 Announce Type: new Abstract: Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves infe

Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card

Model ReleasesDGX agent

arXiv:2604.13466v1 Announce Type: cross Abstract: The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

ApplicationsDGX agent

arXiv:2604.13288v1 Announce Type: new Abstract: We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-t

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

Model ReleasesDGX agent

arXiv:2604.13258v1 Announce Type: new Abstract: Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most exist

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

ResearchDGX agent

arXiv:2604.13977v1 Announce Type: new Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing stra

(How) Learning Rates Regulate Catastrophic Overtraining

ResearchDGX agent

arXiv:2604.13627v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to follow instructions and shaping its behavior as a hel

Hybrid Retrieval for COVID-19 Literature: Comparing Rank Fusion and Projection Fusion with Diversity Reranking

Model ReleasesDGX agent

arXiv:2604.13728v1 Announce Type: cross Abstract: We present a hybrid retrieval system for COVID-19 scientific literature, evaluated on the TREC-COVID benchmark (171,332 papers, 50 expert queries). Th

Indexing Multimodal Language Models for Large-scale Image Retrieval

ResearchDGX agent

arXiv:2604.13268v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remain

IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages

Model ReleasesDGX agent

arXiv:2604.13686v1 Announce Type: new Abstract: While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

Model ReleasesDGX agent

arXiv:2604.13201v1 Announce Type: new Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks

Interpretable Stylistic Variation in Human and LLM Writing Across Genres, Models, and Decoding Strategies

TutorialsDGX agent

arXiv:2604.14111v1 Announce Type: new Abstract: Large Language Models (LLMs) are now capable of generating highly fluent, human-like text. They enable many applications, but also raise concerns such a

IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages

ApplicationsDGX agent

arXiv:2604.13078v1 Announce Type: new Abstract: The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts

Just Use XML: Revisiting Joint Translation and Label Projection

ResearchDGX agent

arXiv:2603.12021v2 Announce Type: replace Abstract: Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from a high-resource language to low-resour

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

Model ReleasesDGX agent

arXiv:2604.13058v1 Announce Type: new Abstract: We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,46

Kwame 2.0: Human-in-the-Loop Generative AI Teaching Assistant for Large Scale Online Coding Education in Africa

TutorialsDGX agent

arXiv:2603.29159v2 Announce Type: replace Abstract: Providing timely and accurate learning support in large-scale online coding courses is challenging, particularly in resource-constrained contexts. W

L2D-Clinical: Learning to Defer for Adaptive Model Selection in Clinical Text Classification

Model ReleasesDGX agent

arXiv:2604.13285v1 Announce Type: new Abstract: Clinical text classification requires choosing between specialized fine-tuned models (BERT variants) and general-purpose large language models (LLMs), y

Language steering in latent space to mitigate unintended code-switching

Model ReleasesDGX agent

arXiv:2510.13849v3 Announce Type: replace Abstract: Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks.

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2511.11334v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian

Learning the Cue or Learning the Word? Analyzing Generalization in Metaphor Detection for Verbs

Model ReleasesDGX agent

arXiv:2604.13713v1 Announce Type: new Abstract: Metaphor detection models achieve strong benchmark performance, yet it remains unclear whether this reflects transferable generalization or lexical memo

Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs

Model ReleasesDGX agent

arXiv:2604.13979v1 Announce Type: new Abstract: Open-world Question Answering (OW-QA) over knowledge graphs (KGs) aims to answer questions over incomplete or evolving KGs. Traditional KGQA assumes a c

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

Model ReleasesDGX agent

arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources

Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning

ApplicationsDGX agent

arXiv:2601.02902v2 Announce Type: replace-cross Abstract: Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decis

Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data

Model ReleasesDGX agent

arXiv:2604.13066v1 Announce Type: new Abstract: In-context learning has established itself as an important learning paradigm for Large Language Models (LLMs). In this paper, we demonstrate that LLMs c

Mathematical Reasoning Enhanced LLM for Formula Derivation: A Case Study on Fiber NLI Modellin

ApplicationsDGX agent

arXiv:2604.13062v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated strong capabilities in code generation and text synthesis, yet their potential for sym

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

Model ReleasesDGX agent

arXiv:2604.13756v1 Announce Type: new Abstract: The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation framework

← Previous
1…116117118119120…128
Next →