AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Agents

How^{2}: How to learn from procedural How-to questions

DGX agent

arXiv:2510.11144v2 Announce Type: replace-cross Abstract: An agent facing a planning problem can use answers to how-to questions to reduce uncertainty and fill knowledge gaps, helping it solve both cu

agentsarxiv-cs-cl
5 May 2026
Model Releases

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2502.16435v4 Announce Type: replace-cross Abstract: Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, cu

model-releasesarxiv-cs-cl
5 May 2026
Safety

Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction

DGX agent

arXiv:2510.25426v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) is positioning language at the core of human-computer interaction (HCI). We argue that advanci

safetyarxiv-cs-cl
5 May 2026
Research

Improving Factuality in LLMs via Inference-Time Knowledge Graph Construction

DGX agent

arXiv:2509.03540v3 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with producing factually consistent answers due to limitations in their parametric memory. Retrieval-Aug

researcharxiv-cs-cl
5 May 2026
Research

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

DGX agent

arXiv:2605.02364v1 Announce Type: new Abstract: Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweigh

researcharxiv-cs-cl
5 May 2026
Safety

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

DGX agent

arXiv:2605.01402v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) struggle with numerical regression under long-tailed target distributions. Token-level supervised fine-tuning (

safetyarxiv-cs-cl
5 May 2026
Research

Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues

DGX agent

arXiv:2605.01097v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the development of AI-powered tutoring systems that provide interactive support via dialogue

researcharxiv-cs-cl
5 May 2026
Safety

IPS: In-Prompt Process Supervision for Short Video Content Moderation

DGX agent

arXiv:2412.15251v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the

safetyarxiv-cs-cl
5 May 2026
Research

Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval

DGX agent

arXiv:2605.02392v1 Announce Type: new Abstract: Novelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an inventio

researcharxiv-cs-cl
5 May 2026
Model Releases

jina-vlm: Small Multilingual Vision Language Model

DGX agent

arXiv:2512.04032v3 Announce Type: replace Abstract: We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation

DGX agent

arXiv:2505.20340v3 Announce Type: replace Abstract: Understanding how latent representations evolve during generation is a central open problem in large language model interpretability. We introduce e

model-releasesarxiv-cs-cl
5 May 2026
Hardware

LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference

DGX agent

arXiv:2605.01058v1 Announce Type: cross Abstract: Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; ye

hardwarearxiv-cs-cl
5 May 2026
Tutorials

Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM

DGX agent

arXiv:2605.01973v1 Announce Type: new Abstract: Conventional LLMs may suffer from corpus heterogeneity and subtle condition changes. While finetuning can create the catastrophe forgetting issue, appli

tutorialsarxiv-cs-cl
5 May 2026
Tutorials

Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative Recommendation

DGX agent

arXiv:2509.10468v2 Announce Type: replace-cross Abstract: Recent advances in generative recommenders adopt a two-stage paradigm: items are first tokenized into semantic IDs using a pretrained tokenize

tutorialsarxiv-cs-cl
5 May 2026
Applications

Led to Mislead: Adversarial Content Injection for Attacks on Neural Ranking Models

DGX agent

arXiv:2605.01591v1 Announce Type: cross Abstract: Neural Ranking Models (NRMs) are central to modern information retrieval but remain highly vulnerable to adversarial manipulation. Existing attacks of

applicationsarxiv-cs-cl
5 May 2026
Safety

Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure

DGX agent

arXiv:2605.01735v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world systems, they must support post-hoc removal of specific content to meet privacy

safetyarxiv-cs-cl
5 May 2026
Research

Leveraging Argument Structure to Predict Content Hatefulness

DGX agent

arXiv:2605.02457v1 Announce Type: new Abstract: Information disorder is a challenging phenomenon that affects society at large. This phenomenon entails the diffusion of misleading, misinforming, and h

researcharxiv-cs-cl
5 May 2026
Research

LITcoder: A General-Purpose Library for Building and Comparing Encoding Models

DGX agent

arXiv:2509.09152v2 Announce Type: replace Abstract: We introduce LITcoder, an open-source library for building and benchmarking neural encoding models. Designed as a flexible backend, LITcoder provide

researcharxiv-cs-cl
5 May 2026
Safety

LLM-Augmented Semantic Steering of Text Embedding Projection Spaces

DGX agent

arXiv:2605.01957v1 Announce Type: cross Abstract: Low-dimensional projections of text embeddings support visual analysis of document collections, but their spatial organization may not reflect the rel

safetyarxiv-cs-cl
5 May 2026
Agents

LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning

DGX agent

arXiv:2605.01047v1 Announce Type: cross Abstract: Hallucinations, outputs that sound plausible but are factually incorrect, remain an open challenge for deployed LLMs. In code generation, models frequ

agentsarxiv-cs-cl
5 May 2026
Hardware

LLM Output Detectability and Task Performance Can be Jointly Optimized

DGX agent

arXiv:2605.01350v1 Announce Type: new Abstract: Detecting machine-generated text is essential for transparency and accountability when deploying large language models (LLMs). Among detection approache

hardwarearxiv-cs-cl
5 May 2026
Safety

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

DGX agent

arXiv:2506.24056v2 Announce Type: replace-cross Abstract: RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce

safetyarxiv-cs-cl
5 May 2026
Research

Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference

DGX agent

arXiv:2411.16821v5 Announce Type: replace Abstract: Non-autoregressive (NAR) language models offer notable efficiency in text generation by circumventing the sequential bottleneck of autoregressive de

researcharxiv-cs-cl
5 May 2026
Agents

Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

DGX agent

arXiv:2605.01224v1 Announce Type: new Abstract: This paper argues that contemporary multilingual NLP has converged on a fragile and misleading paradigm of incidental multilingualism. Today's LLMs appe

agentsarxiv-cs-cl
5 May 2026
Safety

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

DGX agent

arXiv:2605.01347v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single

safetyarxiv-cs-cl
5 May 2026
Hardware

Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models

DGX agent

arXiv:2605.01870v1 Announce Type: new Abstract: Large Language Models (LLMs) have substantially advanced the field of Natural Language Processing (NLP), achieving state-of-the-art performance across a

hardwarearxiv-cs-cl
5 May 2026
Research

Mapping Discourse Reframing: A Multi-Layer Network Approach to Italian HPV Vaccine Discourse on X (2010-2024)

DGX agent

arXiv:2605.02629v1 Announce Type: new Abstract: Understanding how online narratives travel through coalitions is critical for identifying information disorder, yet computational analyses often rely on

researcharxiv-cs-cl
5 May 2026
Research

mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection

DGX agent

arXiv:2605.02712v1 Announce Type: new Abstract: SemEval-2026 Task 10 is focused on conspiracy detection. Specifically, the goal is to detect whether a Reddit comment expresses a conspiracy belief. Our

researcharxiv-cs-cl
5 May 2026
Model Releases

mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection

DGX agent

arXiv:2605.02695v1 Announce Type: new Abstract: SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and mul

model-releasesarxiv-cs-cl
5 May 2026
Tutorials

Measuring AI Reasoning: A Guide for Researchers

DGX agent

arXiv:2605.02442v1 Announce Type: cross Abstract: In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed throug

tutorialsarxiv-cs-cl
5 May 2026
Model Releases

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

DGX agent

arXiv:2605.01417v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insuff

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

DGX agent

arXiv:2605.00969v1 Announce Type: cross Abstract: We present MedMosaic, a medical audio question-answering dataset designed to benchmark language and audio reasoning models under realistic clinical co

model-releasesarxiv-cs-cl
5 May 2026
Research

MemeLens: Multilingual Multitask VLMs for Memes

DGX agent

arXiv:2601.12539v3 Announce Type: replace-cross Abstract: Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery

researcharxiv-cs-cl
5 May 2026
Research

MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents

DGX agent

arXiv:2605.01386v1 Announce Type: new Abstract: Large Language Models (LLMs) lack persistent memory for long-term personalized conversations. Existing graph-based memory systems suffer from informatio

researcharxiv-cs-cl
5 May 2026
Research

Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies

DGX agent

arXiv:2605.02052v1 Announce Type: new Abstract: This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples

researcharxiv-cs-cl
5 May 2026
Research

MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

DGX agent

arXiv:2605.01520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning

researcharxiv-cs-cl
5 May 2026
Safety

Mitigating Misalignment Contagion by Steering with Implicit Traits

DGX agent

arXiv:2605.02751v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are cri

safetyarxiv-cs-cl
5 May 2026
Model Releases

Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models

DGX agent

arXiv:2602.04509v4 Announce Type: replace Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on task-specific data is an effective way to improve performance on downstream applications. Ho

model-releasesarxiv-cs-cl
5 May 2026
Research

Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives

DGX agent

arXiv:2605.00994v1 Announce Type: new Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, rese

researcharxiv-cs-cl
5 May 2026
Applications

Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading

DGX agent

arXiv:2605.01954v1 Announce Type: cross Abstract: Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback i

applicationsarxiv-cs-cl
5 May 2026
Model Releases

MolViBench: Evaluating LLMs on Molecular Vibe Coding

DGX agent

arXiv:2605.02351v1 Announce Type: new Abstract: Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alter

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding

DGX agent

arXiv:2510.08804v3 Announce Type: replace Abstract: We present MOSAIC, a multi-agent Large Language Model (LLM) framework for solving challenging scientific coding tasks. Unlike general-purpose coding

model-releasesarxiv-cs-cl
5 May 2026
Safety

MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation

DGX agent

arXiv:2605.01374v1 Announce Type: new Abstract: Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers

safetyarxiv-cs-cl
5 May 2026
Model Releases

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

DGX agent

arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong

DGX agent

arXiv:2501.09775v3 Announce Type: replace Abstract: Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness o

model-releasesarxiv-cs-cl
5 May 2026
Research

NH-CROP: Robust Pricing for Governed Language Data Assets under Cost Uncertainty

DGX agent

arXiv:2605.01745v1 Announce Type: cross Abstract: Language data are increasingly acquired and governed as assets, yet platforms often price candidate resources before knowing their true privacy or acc

researcharxiv-cs-cl
5 May 2026
Research

Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation

DGX agent

arXiv:2604.03380v2 Announce Type: replace Abstract: Generating diverse, pedagogically valid stories for Arabic early-grade reading assessments requires balancing tight constraints on vocabulary, readi

researcharxiv-cs-cl
5 May 2026
Model Releases

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

DGX agent

arXiv:2605.00877v1 Announce Type: cross Abstract: The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence ha

model-releasesarxiv-cs-cl
5 May 2026
← Previous
1…110111112113114…161
Next →