AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Safety

Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality

DGX agent

arXiv:2603.17512v4 Announce Type: replace Abstract: Large language models (LLMs) exhibit strong general intelligence, yet their multilingual performance remains highly imbalanced. Although LLMs encode

safetyarxiv-cs-cl
17 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

DGX agent

arXiv:2604.14568v1 Announce Type: cross Abstract: Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning.

safetyarxiv-cs-cl
17 Apr 2026
Safety

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

DGX agent

arXiv:2604.15210v1 Announce Type: cross Abstract: Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting the answer right. While recent work evaluates hum

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

DGX agent

arXiv:2512.04578v3 Announce Type: replace Abstract: Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulat

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation

DGX agent

arXiv:2604.14177v1 Announce Type: new Abstract: Grammatical error correction (GEC) and explanation (GEE) have made rapid progress, but real teaching scenarios also require learner-friendly pedagogical

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text

DGX agent

arXiv:2604.14321v1 Announce Type: new Abstract: We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 sur

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

DGX agent

arXiv:2604.14922v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent ad

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

DGX agent

arXiv:2604.15203v1 Announce Type: new Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

DGX agent

arXiv:2604.14448v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select rel

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

MARS^2: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation

DGX agent

arXiv:2604.14564v1 Announce Type: cross Abstract: Reinforcement learning (RL) paradigms have demonstrated strong performance on reasoning-intensive tasks such as code generation. However, limited traj

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

Mechanistic Decoding of Cognitive Constructs in LLMs

DGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation

DGX agent

arXiv:2604.15190v1 Announce Type: cross Abstract: Simulating group-level user behavior enables scalable counterfactual evaluation of merchant strategies without costly online experiments. However, bui

safetyarxiv-cs-cl
17 Apr 2026
Research

MEME-Fusion@CHiPSAL 2026: Multimodal Ablation Study of Hate Detection and Sentiment Analysis on Nepali Memes

DGX agent

arXiv:2604.14218v1 Announce Type: new Abstract: Hate speech detection in Devanagari-scripted social media memes presents compounded challenges: multimodal content structure, script-specific linguistic

researcharxiv-cs-cl
17 Apr 2026
Model Releases

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

DGX agent

arXiv:2604.14158v1 Announce Type: new Abstract: Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Mitigating LLM biases toward spurious social contexts using direct preference optimization

DGX agent

arXiv:2604.02585v2 Announce Type: replace-cross Abstract: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful bia

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

DGX agent

arXiv:2604.14198v1 Announce Type: cross Abstract: Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains large

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

DGX agent

arXiv:2604.15309v1 Announce Type: cross Abstract: The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for we

model-releasesarxiv-cs-cl
17 Apr 2026
Hardware

Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3

DGX agent

arXiv:2603.27844v2 Announce Type: replace Abstract: Majority voting over multiple LLM attempts improves mathematical reasoning, but correlated errors limit the effective sample size. A natural fix is

hardwarearxiv-cs-cl
17 Apr 2026
Safety

Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem

DGX agent

arXiv:2604.14808v1 Announce Type: new Abstract: Machine unlearning for large language models (LLMs) aims to remove targeted knowledge while preserving general capability. In this paper, we recast LLM

safetyarxiv-cs-cl
17 Apr 2026
Safety

Multi-Persona Thinking for Bias Mitigation in Large Language Models

DGX agent

arXiv:2601.15488v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit social biases, which can lead to harmful stereotypes and unfair outcomes. We propose extbf{Multi-Persona Thinki

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

Neuro-Oracle: A Trajectory-Aware Agentic RAG Framework for Interpretable Epilepsy Surgical Prognosis

DGX agent

arXiv:2604.14216v1 Announce Type: cross Abstract: Predicting post-surgical seizure outcomes in pharmacoresistant epilepsy is a clinical challenge. Conventional deep-learning approaches operate on stat

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

NLP needs Diversity outside of 'Diversity'

DGX agent

arXiv:2604.14595v1 Announce Type: new Abstract: This position paper argues that recent progress with diversity in NLP is disproportionately concentrated on a small number of areas surrounding fairness

safetyarxiv-cs-cl
17 Apr 2026
Safety

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

DGX agent

arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a

safetyarxiv-cs-cl
17 Apr 2026
Local Ai

One RL to See Them All: Visual Triple Unified Reinforcement Learning

DGX agent

arXiv:2505.18129v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodolog

local-aiarxiv-cs-cl
17 Apr 2026
Model Releases

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

DGX agent

arXiv:2604.15093v1 Announce Type: cross Abstract: Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achie

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Pangu-ACE: Adaptive Cascaded Experts for Educational Response Generation on EduBench

DGX agent

arXiv:2604.14828v1 Announce Type: new Abstract: Educational assistants should spend more computation only when the task needs it. This paper rewrites our earlier draft around the system that was actua

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

PeerPrism: Peer Evaluation Expertise vs Review-writing AI

DGX agent

arXiv:2604.14513v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, ex

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

POP: Prefill-Only Pruning for Efficient Large Model Inference

DGX agent

arXiv:2602.03295v2 Announce Type: replace Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable capabilities. However, their deployment is hindered by s

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

DGX agent

arXiv:2603.13683v2 Announce Type: replace Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems

DGX agent

arXiv:2604.14585v1 Announce Type: cross Abstract: Prompt optimization in compound AI systems is statistically indistinguishable from a coin flip: across 72 optimization runs on Claude Haiku (6 methods

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking

DGX agent

arXiv:2506.03487v3 Announce Type: replace-cross Abstract: Reranking is fundamental to information retrieval and retrieval-augmented generation, with recent Large Language Models (LLMs) significantly a

model-releasesarxiv-cs-cl
17 Apr 2026
Research

Psychological Steering of Large Language Models

DGX agent

arXiv:2604.14463v1 Announce Type: new Abstract: Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is conver

researcharxiv-cs-cl
17 Apr 2026
Model Releases

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness

DGX agent

arXiv:2604.14324v1 Announce Type: new Abstract: Large language models (LLMs) often exhibit hallucinations due to their inability to accurately perceive their own knowledge boundaries. Existing abstent

model-releasesarxiv-cs-cl
17 Apr 2026
Safety

Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options

DGX agent

arXiv:2604.14634v1 Announce Type: new Abstract: Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by s

safetyarxiv-cs-cl
17 Apr 2026
Safety

QU-NLP at ArchEHR-QA 2026: Two-Stage QLoRA Fine-Tuning of Qwen3-4B for Patient-Oriented Clinical Question Answering and Evidence Sentence Alignment

DGX agent

arXiv:2604.14175v1 Announce Type: new Abstract: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. For

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

DGX agent

arXiv:2604.15151v1 Announce Type: new Abstract: Large language models have demonstrated strong performance on general-purpose programming tasks, yet their ability to generate executable algorithmic tr

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Query pipeline optimization for cancer patient question answering systems

DGX agent

arXiv:2412.14751v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) mitigates hallucination in Large Language Models (LLMs) by using query pipelines to retrieve relevant external

model-releasesarxiv-cs-cl
17 Apr 2026
Research

RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding

DGX agent

arXiv:2604.14885v1 Announce Type: new Abstract: Autoregressive decoding in Large Language Models (LLMs) generates one token per step, causing high inference latency. Speculative decoding (SD) mitigate

researcharxiv-cs-cl
17 Apr 2026
Safety

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

DGX agent

arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,

safetyarxiv-cs-cl
17 Apr 2026
Safety

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

DGX agent

arXiv:2604.14888v1 Announce Type: new Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains

safetyarxiv-cs-cl
17 Apr 2026
Research

ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation

DGX agent

arXiv:2603.21094v3 Announce Type: replace Abstract: Human annotation is central to NLP evaluation, yet subjective tasks often exhibit substantial variability across annotators. While large language mo

researcharxiv-cs-cl
17 Apr 2026
Model Releases

Rethinking Patient Education as Multi-turn Multi-modal Interaction

DGX agent

arXiv:2604.14656v1 Announce Type: cross Abstract: Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient ed

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring

DGX agent

arXiv:2604.14616v1 Announce Type: new Abstract: Clinical value set authoring -- the task of identifying all codes in a standardized vocabulary that define a clinical concept -- is a recurring bottlene

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

DGX agent

arXiv:2604.14261v1 Announce Type: new Abstract: The rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support. However, LLM-base

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification

DGX agent

arXiv:2604.05302v2 Announce Type: replace Abstract: Text simplification supports second language (L2) learning by providing comprehensible input, consistent with the Input Hypothesis. However, constru

model-releasesarxiv-cs-cl
17 Apr 2026
Applications

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

DGX agent

arXiv:2604.15022v1 Announce Type: cross Abstract: Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing

applicationsarxiv-cs-cl
17 Apr 2026
Model Releases

SAGE Celer 2.6 Technical Card

DGX agent

arXiv:2604.14168v1 Announce Type: new Abstract: We introduce SAGE Celer 2.6, the latest in our line of general-purpose Celer models from SAGEA. Celer 2.6 is available in 5B, 10B, and 27B parameter siz

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding

DGX agent

arXiv:2604.14862v1 Announce Type: new Abstract: Constrained decoding has been widely adopted for structured generation with large language models (LLMs), ensuring that outputs satisfy predefined forma

model-releasesarxiv-cs-cl
17 Apr 2026
← Previous
1…143144145146147…160
Next →