AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Model Releases

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

DGX agent

arXiv:2603.02097v5 Announce Type: replace Abstract: Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in l

model-releasesarxiv-cs-cl
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Local Ai

ClinicalEncoder26AM: A Multlilingual Diagnosable ColBERT Model; Evidences from the MultiClinNER Shared Task

DGX agent

arXiv:2605.28521v1 Announce Type: new Abstract: ClinicalEncoder26AM is a multilingual Diagnosable ColBERT for clinical and biomedical texts, which aligns at multiple levels its token-level semantic wi

local-aiarxiv-cs-cl
28 May 2026
Model Releases

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

DGX agent

arXiv:2605.28734v1 Announce Type: cross Abstract: A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a work

model-releasesarxiv-cs-cl
28 May 2026
Safety

CodeGENCAT: Generative Computerized Adaptive Testing for Open-ended Coding Problems

DGX agent

arXiv:2602.20020v2 Announce Type: replace Abstract: Existing Computerized Adaptive Testing (CAT) frameworks typically select questions based on the predicted likelihood that the student will answer co

safetyarxiv-cs-cl
28 May 2026
Research

Comonadic Morphophonology: A Compositional Framework for Context-Dependent Morphological Rules in Finnish

DGX agent

arXiv:2605.28484v1 Announce Type: new Abstract: Composing finite-state transducers (FSTs) for context-dependent morphophonological rules -- consonant gradation, vowel harmony, possessive suffix assimi

researcharxiv-cs-cl
28 May 2026
Model Releases

ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

DGX agent

arXiv:2605.28093v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA)

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor

DGX agent

arXiv:2605.28062v1 Announce Type: new Abstract: We describe ConvMemory, a small 3.6M-parameter learned reranker for conversational long-term memory retrieval, trained with cross-encoder teacher superv

model-releasesarxiv-cs-cl
28 May 2026
Local Ai

Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance

DGX agent

arXiv:2602.03491v2 Announce Type: replace-cross Abstract: Reasoning over table images remains challenging for Large Vision-Language Models (LVLMs) due to complex layouts and tightly coupled structure-

local-aiarxiv-cs-cl
28 May 2026
Model Releases

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

DGX agent

arXiv:2605.27957v1 Announce Type: new Abstract: Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Disentangling Language Roles in Multilingual LLM Task Execution

DGX agent

arXiv:2605.27649v1 Announce Type: new Abstract: Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expa

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation

DGX agent

arXiv:2603.21465v2 Announce Type: replace Abstract: Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Mode

model-releasesarxiv-cs-cl
28 May 2026
Safety

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

DGX agent

arXiv:2605.27741v1 Announce Type: new Abstract: Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post

safetyarxiv-cs-cl
28 May 2026
Research

Evaluating the Generation Capabilities of Large Chinese Language Models

DGX agent

arXiv:2308.04823v5 Announce Type: replace Abstract: This paper unveils CG-Eval, the first-ever comprehensive and automated evaluation framework designed for assessing the generative capabilities of la

researcharxiv-cs-cl
28 May 2026
Research

Explanation Generation for Contradiction Reconciliation with LLMs

DGX agent

arXiv:2603.22735v2 Announce Type: replace Abstract: Existing NLP work commonly treats contradictions as errors to be resolved by choosing which statements to accept or discard. Yet a key aspect of hum

researcharxiv-cs-cl
28 May 2026
Safety

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

DGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

safetyarxiv-cs-cl
28 May 2026
Research

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation

DGX agent

arXiv:2601.03549v2 Announce Type: replace-cross Abstract: Sign Language Translation (SLT) is a challenging cross-modal task requiring joint modeling of manual articulations and non-manual signals. Exi

researcharxiv-cs-cl
28 May 2026
Applications

FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations

DGX agent

arXiv:2605.27896v1 Announce Type: new Abstract: Recently, large language models (LLMs) have achieved superior performance in static financial reasoning and simple dynamic trading tasks. However, exist

applicationsarxiv-cs-cl
28 May 2026
Research

Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models

DGX agent

arXiv:2510.17620v2 Announce Type: replace Abstract: Large language models may encode sensitive information or outdated knowledge that needs to be removed, to ensure responsible and compliant model res

researcharxiv-cs-cl
28 May 2026
Tutorials

Formula-One Prompting: A Composable Equation-First Prefix for Applied Mathematics

DGX agent

arXiv:2601.19302v3 Announce Type: replace Abstract: This paper introduces Formula Prompting (FP) and Formula-One Prompting (F-1), two single-call methods that elicit governing equations before solving

tutorialsarxiv-cs-cl
28 May 2026
Model Releases

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

DGX agent

arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall

model-releasesarxiv-cs-cl
28 May 2026
Safety

GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization

DGX agent

arXiv:2605.27934v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards,

safetyarxiv-cs-cl
28 May 2026
Research

GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors

DGX agent

arXiv:2605.27866v1 Announce Type: new Abstract: Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionab

researcharxiv-cs-cl
28 May 2026
Research

GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study

DGX agent

arXiv:2605.28643v1 Announce Type: new Abstract: Methods to represent literary texts as graphs or sequences of graphs mainly focus on representing character interactions, and often overlook another cru

researcharxiv-cs-cl
28 May 2026
Applications

GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction

DGX agent

arXiv:2605.28645v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Grap

applicationsarxiv-cs-cl
28 May 2026
Agents

GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection

DGX agent

arXiv:2605.28534v1 Announce Type: new Abstract: Despite the rapid progress of multimodal large language models in building Graphical User Interface (GUI) agents, their real-world task completion is fu

agentsarxiv-cs-cl
28 May 2026
Model Releases

HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

DGX agent

arXiv:2605.28315v1 Announce Type: new Abstract: General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

DGX agent

arXiv:2605.28308v1 Announce Type: new Abstract: Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relat

model-releasesarxiv-cs-cl
28 May 2026
Tutorials

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

DGX agent

arXiv:2605.28802v1 Announce Type: new Abstract: Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisi

tutorialsarxiv-cs-cl
28 May 2026
Safety

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

DGX agent

arXiv:2605.27374v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, pers

safetyarxiv-cs-cl
28 May 2026
Safety

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

DGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

safetyarxiv-cs-cl
28 May 2026
Model Releases

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following

DGX agent

arXiv:2605.28218v1 Announce Type: new Abstract: Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated gloss

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

DGX agent

arXiv:2605.28649v1 Announce Type: cross Abstract: LLMs increasingly require surgical model editing to enhance domain-specific capabilities without incurring the computational cost or catastrophic forg

model-releasesarxiv-cs-cl
28 May 2026
Tutorials

Keyphrase Generative Representation of Youth Crisis Conversations Beyond Static Taxonomies

DGX agent

arXiv:2605.27546v1 Announce Type: new Abstract: Crisis Responders (CRs) rapidly assess thousands of youth SMS conversations each year to identify mental health concerns and guide support. Yet youth di

tutorialsarxiv-cs-cl
28 May 2026
Agents

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

DGX agent

arXiv:2605.27788v1 Announce Type: cross Abstract: Humans know when to reach for help e.g. 347 imes 28 warrants a calculator while 2+2 does not. Language models do not. Prompt-based approaches can inst

agentsarxiv-cs-cl
28 May 2026
Research

Knowledge Dependency Estimation for Reliable Question Answering

DGX agent

arXiv:2605.28047v1 Announce Type: new Abstract: Reliable question answering requires identifying not only whether an answer is correct, but also which available knowledge the prediction depends on. In

researcharxiv-cs-cl
28 May 2026
Model Releases

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

DGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

model-releasesarxiv-cs-cl
28 May 2026
Agents

Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles

DGX agent

arXiv:2512.20780v3 Announce Type: replace Abstract: Recent work has explored the use of large language models (LLMs) to generate tutoring responses in mathematics, yet it remains unclear how closely t

agentsarxiv-cs-cl
28 May 2026
Agents

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis

DGX agent

arXiv:2601.16800v3 Announce Type: replace Abstract: Fine-grained opinion analysis of text provides a detailed understanding of expressed sentiments, including the addressed entity. Although this level

agentsarxiv-cs-cl
28 May 2026
Model Releases

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

DGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

model-releasesarxiv-cs-cl
28 May 2026
Safety

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

DGX agent

arXiv:2507.06999v2 Announce Type: replace-cross Abstract: Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal

safetyarxiv-cs-cl
28 May 2026
Model Releases

Learning to Translate from Soft to Hard LLM Prompts

DGX agent

arXiv:2605.27642v1 Announce Type: new Abstract: Soft prompt tuning is a parameter-efficient method for adapting LLMs to specific tasks, but suffers from a lack of interpretability. Building on recent

model-releasesarxiv-cs-cl
28 May 2026
Hardware

Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems

DGX agent

arXiv:2605.27787v1 Announce Type: cross Abstract: Multi-agent systems (MAS) have substantially advanced autonomous software engineering (SWE), but their growing inference energy demands raise sustaina

hardwarearxiv-cs-cl
28 May 2026
Model Releases

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

DGX agent

arXiv:2603.21165v2 Announce Type: replace Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresente

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

MaskClaw: Edge-Side Personalized Privacy Arbitration for GUI Agents with Behavior-Driven Skill Evolution

DGX agent

arXiv:2605.28646v1 Announce Type: cross Abstract: GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records,

model-releasesarxiv-cs-cl
28 May 2026
Research

MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment

DGX agent

arXiv:2605.27865v1 Announce Type: new Abstract: Matching submissions with suitable reviewers at scale is a growing challenge for major venues, yet existing approaches either rely on coarse proxy signa

researcharxiv-cs-cl
28 May 2026
Safety

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

DGX agent

arXiv:2605.12515v2 Announce Type: replace Abstract: Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's langua

safetyarxiv-cs-cl
28 May 2026
Safety

Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

DGX agent

arXiv:2605.28629v1 Announce Type: new Abstract: Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute

safetyarxiv-cs-cl
28 May 2026
← Previous
1…7172737475…162
Next →