AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 May 2026

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM

TutorialsDGX agent

arXiv:2605.01973v1 Announce Type: new Abstract: Conventional LLMs may suffer from corpus heterogeneity and subtle condition changes. While finetuning can create the catastrophe forgetting issue, appli

Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative Recommendation

TutorialsDGX agent

arXiv:2509.10468v2 Announce Type: replace-cross Abstract: Recent advances in generative recommenders adopt a two-stage paradigm: items are first tokenized into semantic IDs using a pretrained tokenize

Led to Mislead: Adversarial Content Injection for Attacks on Neural Ranking Models

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.01591v1 Announce Type: cross Abstract: Neural Ranking Models (NRMs) are central to modern information retrieval but remain highly vulnerable to adversarial manipulation. Existing attacks of

Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure

SafetyDGX agent

arXiv:2605.01735v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world systems, they must support post-hoc removal of specific content to meet privacy

Leveraging Argument Structure to Predict Content Hatefulness

ResearchDGX agent

arXiv:2605.02457v1 Announce Type: new Abstract: Information disorder is a challenging phenomenon that affects society at large. This phenomenon entails the diffusion of misleading, misinforming, and h

LITcoder: A General-Purpose Library for Building and Comparing Encoding Models

ResearchDGX agent

arXiv:2509.09152v2 Announce Type: replace Abstract: We introduce LITcoder, an open-source library for building and benchmarking neural encoding models. Designed as a flexible backend, LITcoder provide

LLM-Augmented Semantic Steering of Text Embedding Projection Spaces

SafetyDGX agent

arXiv:2605.01957v1 Announce Type: cross Abstract: Low-dimensional projections of text embeddings support visual analysis of document collections, but their spatial organization may not reflect the rel

LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning

AgentsDGX agent

arXiv:2605.01047v1 Announce Type: cross Abstract: Hallucinations, outputs that sound plausible but are factually incorrect, remain an open challenge for deployed LLMs. In code generation, models frequ

LLM Output Detectability and Task Performance Can be Jointly Optimized

HardwareDGX agent

arXiv:2605.01350v1 Announce Type: new Abstract: Detecting machine-generated text is essential for transparency and accountability when deploying large language models (LLMs). Among detection approache

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

SafetyDGX agent

arXiv:2506.24056v2 Announce Type: replace-cross Abstract: RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce

Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference

ResearchDGX agent

arXiv:2411.16821v5 Announce Type: replace Abstract: Non-autoregressive (NAR) language models offer notable efficiency in text generation by circumventing the sequential bottleneck of autoregressive de

Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

AgentsDGX agent

arXiv:2605.01224v1 Announce Type: new Abstract: This paper argues that contemporary multilingual NLP has converged on a fragile and misleading paradigm of incidental multilingualism. Today's LLMs appe

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

SafetyDGX agent

arXiv:2605.01347v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single

Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models

HardwareDGX agent

arXiv:2605.01870v1 Announce Type: new Abstract: Large Language Models (LLMs) have substantially advanced the field of Natural Language Processing (NLP), achieving state-of-the-art performance across a

Mapping Discourse Reframing: A Multi-Layer Network Approach to Italian HPV Vaccine Discourse on X (2010-2024)

ResearchDGX agent

arXiv:2605.02629v1 Announce Type: new Abstract: Understanding how online narratives travel through coalitions is critical for identifying information disorder, yet computational analyses often rely on

mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection

ResearchDGX agent

arXiv:2605.02712v1 Announce Type: new Abstract: SemEval-2026 Task 10 is focused on conspiracy detection. Specifically, the goal is to detect whether a Reddit comment expresses a conspiracy belief. Our

mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection

Model ReleasesDGX agent

arXiv:2605.02695v1 Announce Type: new Abstract: SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and mul

Measuring AI Reasoning: A Guide for Researchers

TutorialsDGX agent

arXiv:2605.02442v1 Announce Type: cross Abstract: In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed throug

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

Model ReleasesDGX agent

arXiv:2605.01417v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insuff

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

Model ReleasesDGX agent

arXiv:2605.00969v1 Announce Type: cross Abstract: We present MedMosaic, a medical audio question-answering dataset designed to benchmark language and audio reasoning models under realistic clinical co

MemeLens: Multilingual Multitask VLMs for Memes

ResearchDGX agent

arXiv:2601.12539v3 Announce Type: replace-cross Abstract: Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery

MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents

ResearchDGX agent

arXiv:2605.01386v1 Announce Type: new Abstract: Large Language Models (LLMs) lack persistent memory for long-term personalized conversations. Existing graph-based memory systems suffer from informatio

Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies

ResearchDGX agent

arXiv:2605.02052v1 Announce Type: new Abstract: This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples

MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

ResearchDGX agent

arXiv:2605.01520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning

Mitigating Misalignment Contagion by Steering with Implicit Traits

SafetyDGX agent

arXiv:2605.02751v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are cri

Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2602.04509v4 Announce Type: replace Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on task-specific data is an effective way to improve performance on downstream applications. Ho

Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives

ResearchDGX agent

arXiv:2605.00994v1 Announce Type: new Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, rese

Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading

ApplicationsDGX agent

arXiv:2605.01954v1 Announce Type: cross Abstract: Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback i

MolViBench: Evaluating LLMs on Molecular Vibe Coding

Model ReleasesDGX agent

arXiv:2605.02351v1 Announce Type: new Abstract: Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alter

MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding

Model ReleasesDGX agent

arXiv:2510.08804v3 Announce Type: replace Abstract: We present MOSAIC, a multi-agent Large Language Model (LLM) framework for solving challenging scientific coding tasks. Unlike general-purpose coding

MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation

SafetyDGX agent

arXiv:2605.01374v1 Announce Type: new Abstract: Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

Model ReleasesDGX agent

arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong

Model ReleasesDGX agent

arXiv:2501.09775v3 Announce Type: replace Abstract: Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness o

NH-CROP: Robust Pricing for Governed Language Data Assets under Cost Uncertainty

ResearchDGX agent

arXiv:2605.01745v1 Announce Type: cross Abstract: Language data are increasingly acquired and governed as assets, yet platforms often price candidate resources before knowing their true privacy or acc

Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation

ResearchDGX agent

arXiv:2604.03380v2 Announce Type: replace Abstract: Generating diverse, pedagogically valid stories for Arabic early-grade reading assessments requires balancing tight constraints on vocabulary, readi

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

Model ReleasesDGX agent

arXiv:2605.00877v1 Announce Type: cross Abstract: The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence ha

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

Model ReleasesDGX agent

arXiv:2510.07731v3 Announce Type: replace-cross Abstract: Organic reaction mechanisms are the stepwise elementary reactions by which reactants form intermediates and products, and are fundamental to u

On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility

Model ReleasesDGX agent

arXiv:2605.01357v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily

On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval

Model ReleasesDGX agent

arXiv:2506.11499v2 Announce Type: replace Abstract: Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have she

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

ResearchDGX agent

arXiv:2605.01749v1 Announce Type: new Abstract: Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors

OpenAI GPT-5 System Card

Model ReleasesDGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

Model ReleasesDGX agent

arXiv:2605.01333v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-lev

Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models

Model ReleasesDGX agent

arXiv:2511.21086v2 Announce Type: replace Abstract: Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains l

Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring

Model ReleasesDGX agent

arXiv:2605.02069v1 Announce Type: new Abstract: Many scoring applications require absolute predictions, while pairwise comparisons can provide a simpler learning objective. We present Pair2Score, a tw

PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention

Model ReleasesDGX agent

arXiv:2605.02447v1 Announce Type: new Abstract: Multimodal sarcasm detection, which aims to precisely identify pragmatic incongruities between literal text and nonverbal cues, has gained substantial a

Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates

Model ReleasesDGX agent

arXiv:2605.02236v1 Announce Type: cross Abstract: Recursive language-model loops often settle into recognizable attractor-like patterns. The practical question is how much injected text is needed to m

Prescriptive Scaling Laws for Data Constrained Training

Model ReleasesDGX agent

arXiv:2605.01640v1 Announce Type: cross Abstract: Training compute is increasingly outpacing the availability of high-quality data. This shifts the central challenge from optimal compute allocation to

Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm

Model ReleasesDGX agent

arXiv:2602.11543v2 Announce Type: replace Abstract: Pretraining large language models (LLMs) typically requires centralized clusters with thousands of high-memory GPUs (e.g., H100/A100). Recent decent

Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation

ResearchDGX agent

arXiv:2512.06938v2 Announce Type: replace Abstract: Modern neural language models achieve high accuracy in text generation, yet precise control over generation length remains underdeveloped. In this p

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

Psychologically Potent, Computationally Invisible: LLMs Generate Social-Comparison Triggers They Fail to Detect

Model ReleasesDGX agent

arXiv:2605.01017v1 Announce Type: new Abstract: We introduce Xiaohongshu Social Comparison Reader Elicitation (XHS-SCoRE), a reader-grounded benchmark for detecting if a text-only Xiaohongshu (RedNote

PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature

ResearchDGX agent

arXiv:2605.02720v1 Announce Type: cross Abstract: Vision-language models hold considerable promise for ophthalmology, but their development depends on large-scale, high-quality image-text datasets tha

Quantifying and Predicting Disagreement in Graded Human Ratings

ResearchDGX agent

arXiv:2605.01168v1 Announce Type: new Abstract: It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all in

RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions

TutorialsDGX agent

arXiv:2605.01104v1 Announce Type: cross Abstract: Understanding how developers interact with AI coding assistants requires more than chat logs or git histories in isolation; it requires reconstructing

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

SafetyDGX agent

arXiv:2509.10746v3 Announce Type: replace Abstract: Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for cl

ReFRAME or Remain: Unsupervised Lexical Semantic Change Detection with Frame Semantics

ResearchDGX agent

arXiv:2602.04514v3 Announce Type: replace Abstract: The majority of contemporary computational methods for lexical semantic change (LSC) detection are based on neural embedding distributional represen

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

Model ReleasesDGX agent

arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces

Model ReleasesDGX agent

arXiv:2605.02801v1 Announce Type: new Abstract: As large language model (LLM) agents evolve from isolated tool users into coordinated teams, reinforcement learning (RL) must optimize not only individu

Reinforcement Learning from Compiler and Language Server Feedback

SafetyDGX agent

arXiv:2510.22907v2 Announce Type: replace Abstract: Coding agents fail when text-level guesses outrun program facts: they hallucinate APIs, drift to the wrong symbol, and apply edits without evidence

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

SafetyDGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

← Previous
1…8889909192…129
Next →