AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
17 Apr 2026

IG-Search: Step-Level Information Gain Rewards for Search-Augmented Reasoning

SafetyDGX agent

arXiv:2604.15148v1 Announce Type: cross Abstract: Reinforcement learning has emerged as an effective paradigm for training large language models to perform search-augmented reasoning. However, existin

Improving Language Models with Intentional Analysis

Model ReleasesDGX agent

arXiv:2502.04689v4 Announce Type: replace Abstract: Intent, a critical cognitive notion and mental state, is ubiquitous in human communication and problem-solving. Accurately understanding the underly

In Context Learning and Reasoning for Symbolic Regression with Large Language Models

Model ReleasesDGX agent

arXiv:2410.17448v3 Announce Type: replace Abstract: Large Language Models (LLMs) are transformer-based machine learning models that have shown remarkable performance in tasks for which they were not e


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model

Model ReleasesDGX agent

arXiv:2604.14180v1 Announce Type: new Abstract: We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, with zero Englis

IROSA: Interactive Robot Skill Adaptation using Natural Language

SafetyDGX agent

arXiv:2603.03897v3 Announce Type: replace-cross Abstract: Foundation models have demonstrated impressive capabilities across diverse domains, while imitation learning provides principled methods for r

IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation

ApplicationsDGX agent

arXiv:2604.15109v1 Announce Type: new Abstract: Despite the rapid advancement of Large Language Models (LLMs), uncertainty quantification in LLM generation is a persistent challenge. Although recent a

Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER

ResearchDGX agent

arXiv:2604.05158v2 Announce Type: replace Abstract: Large language models encode extensive world knowledge valuable for zero-shot named entity recognition. However, their causal attention mechanism, w

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

Model ReleasesDGX agent

arXiv:2604.14799v1 Announce Type: new Abstract: Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing e

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality

TutorialsDGX agent

arXiv:2506.19807v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs), particularly slow-thinking models, often exhibit severe hallucination, outputting incorrect content due to an in

Language Model as Planner and Formalizer under Constraints

SafetyDGX agent

arXiv:2510.05486v2 Announce Type: replace Abstract: LLMs have been widely used in planning, either as planners to generate action sequences end-to-end, or as formalizers to represent the planning doma

Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions

ApplicationsDGX agent

arXiv:2502.16761v2 Announce Type: replace Abstract: Large language models (LLMs) present novel opportunities in public opinion research by predicting survey responses in advance during the early stage

Language of Thought Shapes Output Diversity in Large Language Models

SafetyDGX agent

arXiv:2601.11227v2 Announce Type: replace Abstract: Output diversity is crucial for Large Language Models as it underpins pluralism and creativity. In this work, we reveal that controlling the languag

Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality

SafetyDGX agent

arXiv:2603.17512v4 Announce Type: replace Abstract: Large language models (LLMs) exhibit strong general intelligence, yet their multilingual performance remains highly imbalanced. Although LLMs encode

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

SafetyDGX agent

arXiv:2604.14568v1 Announce Type: cross Abstract: Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning.

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

SafetyDGX agent

arXiv:2604.15210v1 Announce Type: cross Abstract: Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting the answer right. While recent work evaluates hum

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

Model ReleasesDGX agent

arXiv:2512.04578v3 Announce Type: replace Abstract: Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulat

Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation

Model ReleasesDGX agent

arXiv:2604.14177v1 Announce Type: new Abstract: Grammatical error correction (GEC) and explanation (GEE) have made rapid progress, but real teaching scenarios also require learner-friendly pedagogical

LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text

Model ReleasesDGX agent

arXiv:2604.14321v1 Announce Type: new Abstract: We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 sur

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.14922v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent ad

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

Model ReleasesDGX agent

arXiv:2604.15203v1 Announce Type: new Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

Model ReleasesDGX agent

arXiv:2604.14448v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select rel

MARS^2: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation

SafetyDGX agent

arXiv:2604.14564v1 Announce Type: cross Abstract: Reinforcement learning (RL) paradigms have demonstrated strong performance on reasoning-intensive tasks such as code generation. However, limited traj

Mechanistic Decoding of Cognitive Constructs in LLMs

Model ReleasesDGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation

SafetyDGX agent

arXiv:2604.15190v1 Announce Type: cross Abstract: Simulating group-level user behavior enables scalable counterfactual evaluation of merchant strategies without costly online experiments. However, bui

MEME-Fusion@CHiPSAL 2026: Multimodal Ablation Study of Hate Detection and Sentiment Analysis on Nepali Memes

ResearchDGX agent

arXiv:2604.14218v1 Announce Type: new Abstract: Hate speech detection in Devanagari-scripted social media memes presents compounded challenges: multimodal content structure, script-specific linguistic

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

Model ReleasesDGX agent

arXiv:2604.14158v1 Announce Type: new Abstract: Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the

Mitigating LLM biases toward spurious social contexts using direct preference optimization

Model ReleasesDGX agent

arXiv:2604.02585v2 Announce Type: replace-cross Abstract: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful bia

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

Model ReleasesDGX agent

arXiv:2604.14198v1 Announce Type: cross Abstract: Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains large

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

Model ReleasesDGX agent

arXiv:2604.15309v1 Announce Type: cross Abstract: The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for we

Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3

HardwareDGX agent

arXiv:2603.27844v2 Announce Type: replace Abstract: Majority voting over multiple LLM attempts improves mathematical reasoning, but correlated errors limit the effective sample size. A natural fix is

Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem

SafetyDGX agent

arXiv:2604.14808v1 Announce Type: new Abstract: Machine unlearning for large language models (LLMs) aims to remove targeted knowledge while preserving general capability. In this paper, we recast LLM

Multi-Persona Thinking for Bias Mitigation in Large Language Models

SafetyDGX agent

arXiv:2601.15488v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose extbf{Multi-Persona Thinki

Neuro-Oracle: A Trajectory-Aware Agentic RAG Framework for Interpretable Epilepsy Surgical Prognosis

Model ReleasesDGX agent

arXiv:2604.14216v1 Announce Type: cross Abstract: Predicting post-surgical seizure outcomes in pharmacoresistant epilepsy is a clinical challenge. Conventional deep-learning approaches operate on stat

NLP needs Diversity outside of 'Diversity'

SafetyDGX agent

arXiv:2604.14595v1 Announce Type: new Abstract: This position paper argues that recent progress with diversity in NLP is disproportionately concentrated on a small number of areas surrounding fairness

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

SafetyDGX agent

arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Local AiDGX agent

arXiv:2505.18129v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodolog

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Model ReleasesDGX agent

arXiv:2604.15093v1 Announce Type: cross Abstract: Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achie

Pangu-ACE: Adaptive Cascaded Experts for Educational Response Generation on EduBench

Model ReleasesDGX agent

arXiv:2604.14828v1 Announce Type: new Abstract: Educational assistants should spend more computation only when the task needs it. This paper rewrites our earlier draft around the system that was actua

PeerPrism: Peer Evaluation Expertise vs Review-writing AI

Model ReleasesDGX agent

arXiv:2604.14513v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, ex

POP: Prefill-Only Pruning for Efficient Large Model Inference

Model ReleasesDGX agent

arXiv:2602.03295v2 Announce Type: replace Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable capabilities. However, their deployment is hindered by s

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

SafetyDGX agent

arXiv:2603.13683v2 Announce Type: replace Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems

Model ReleasesDGX agent

arXiv:2604.14585v1 Announce Type: cross Abstract: Prompt optimization in compound AI systems is statistically indistinguishable from a coin flip: across 72 optimization runs on Claude Haiku (6 methods

ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking

Model ReleasesDGX agent

arXiv:2506.03487v3 Announce Type: replace-cross Abstract: Reranking is fundamental to information retrieval and retrieval-augmented generation, with recent Large Language Models (LLMs) significantly a

Psychological Steering of Large Language Models

ResearchDGX agent

arXiv:2604.14463v1 Announce Type: new Abstract: Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is conver

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness

Model ReleasesDGX agent

arXiv:2604.14324v1 Announce Type: new Abstract: Large language models (LLMs) often exhibit hallucinations due to their inability to accurately perceive their own knowledge boundaries. Existing abstent

Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options

SafetyDGX agent

arXiv:2604.14634v1 Announce Type: new Abstract: Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by s

QU-NLP at ArchEHR-QA 2026: Two-Stage QLoRA Fine-Tuning of Qwen3-4B for Patient-Oriented Clinical Question Answering and Evidence Sentence Alignment

SafetyDGX agent

arXiv:2604.14175v1 Announce Type: new Abstract: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. For

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

Model ReleasesDGX agent

arXiv:2604.15151v1 Announce Type: new Abstract: Large language models have demonstrated strong performance on general-purpose programming tasks, yet their ability to generate executable algorithmic tr

Query pipeline optimization for cancer patient question answering systems

Model ReleasesDGX agent

arXiv:2412.14751v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) mitigates hallucination in Large Language Models (LLMs) by using query pipelines to retrieve relevant external

RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding

ResearchDGX agent

arXiv:2604.14885v1 Announce Type: new Abstract: Autoregressive decoding in Large Language Models (LLMs) generates one token per step, causing high inference latency. Speculative decoding (SD) mitigate

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

SafetyDGX agent

arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

SafetyDGX agent

arXiv:2604.14888v1 Announce Type: new Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains

ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation

ResearchDGX agent

arXiv:2603.21094v3 Announce Type: replace Abstract: Human annotation is central to NLP evaluation, yet subjective tasks often exhibit substantial variability across annotators. While large language mo

Rethinking Patient Education as Multi-turn Multi-modal Interaction

Model ReleasesDGX agent

arXiv:2604.14656v1 Announce Type: cross Abstract: Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient ed

Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring

Model ReleasesDGX agent

arXiv:2604.14616v1 Announce Type: new Abstract: Clinical value set authoring -- the task of identifying all codes in a standardized vocabulary that define a clinical concept -- is a recurring bottlene

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

Model ReleasesDGX agent

arXiv:2604.14261v1 Announce Type: new Abstract: The rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support. However, LLM-base

Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification

Model ReleasesDGX agent

arXiv:2604.05302v2 Announce Type: replace Abstract: Text simplification supports second language (L2) learning by providing comprehensible input, consistent with the Input Hypothesis. However, constru

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

ApplicationsDGX agent

arXiv:2604.15022v1 Announce Type: cross Abstract: Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing

SAGE Celer 2.6 Technical Card

Model ReleasesDGX agent

arXiv:2604.14168v1 Announce Type: new Abstract: We introduce SAGE Celer 2.6, the latest in our line of general-purpose Celer models from SAGEA. Celer 2.6 is available in 5B, 10B, and 27B parameter siz

Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding

Model ReleasesDGX agent

arXiv:2604.14862v1 Announce Type: new Abstract: Constrained decoding has been widely adopted for structured generation with large language models (LLMs), ensuring that outputs satisfy predefined forma

← Previous
1…114115116117118…128
Next →