AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines

DGX agent

arXiv:2601.03627v3 Announce Type: replace Abstract: We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs

model-releasesarxiv-cs-cl
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs

DGX agent

arXiv:2505.02072v2 Announce Type: replace Abstract: Language modeling has shifted in recent years from a distribution over strings to prediction models with textual inputs and outputs for general-purp

researcharxiv-cs-cl
13 May 2026
Model Releases

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum

DGX agent

arXiv:2605.11403v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Opti

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

DGX agent

arXiv:2504.14707v3 Announce Type: replace Abstract: We introduce FLAME (FLemish Accounts of Momentary Experiences), a new corpus of nearly 25,000 daily personal narratives in Belgian-Dutch (Flemish),

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

DGX agent

arXiv:2605.11416v1 Announce Type: new Abstract: Selective layer-wise updates are essential for low-cost continued pre-training of Large Language Models (LLMs), yet determining which layers to freeze o

model-releasesarxiv-cs-cl
13 May 2026
Applications

From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction

DGX agent

arXiv:2605.11774v1 Announce Type: new Abstract: By processing electronic health records (EHRs) as natural language sequences, large language models (LLMs) have shown potential in clinical prediction t

applicationsarxiv-cs-cl
13 May 2026
Safety

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

DGX agent

arXiv:2605.11853v1 Announce Type: cross Abstract: Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that pr

safetyarxiv-cs-cl
13 May 2026
Model Releases

Geometric Factual Recall in Transformers

DGX agent

arXiv:2605.12426v1 Announce Type: new Abstract: How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of emb

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender

DGX agent

arXiv:2605.12299v1 Announce Type: new Abstract: Recent works have analyzed the impact of individual components of neural networks on gendered predictions, often with a focus on mitigating gender bias.

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction

DGX agent

arXiv:2305.12678v3 Announce Type: replace Abstract: Multimodal Review Helpfulness Prediction (MRHP) aims to rank product reviews based on predicted helpfulness scores and has been widely applied in e-

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

DGX agent

arXiv:2605.06152v2 Announce Type: replace-cross Abstract: Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the 'Slingshot Mechanism.' Ex

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs

DGX agent

arXiv:2503.06139v2 Announce Type: replace Abstract: Pairwise LLM-as-a-judge evaluation asks the judge to identify the better of two candidate answers. We study a one-line modification that asks for th

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

DGX agent

arXiv:2601.20255v2 Announce Type: replace-cross Abstract: SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabil

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model

DGX agent

arXiv:2605.11255v1 Announce Type: new Abstract: We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Traini

model-releasesarxiv-cs-cl
13 May 2026
Applications

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation

DGX agent

arXiv:2605.11651v1 Announce Type: cross Abstract: Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the fi

applicationsarxiv-cs-cl
13 May 2026
Safety

How Does Differential Privacy Affect Social Bias in LLMs? A Systematic Evaluation

DGX agent

arXiv:2605.11195v1 Announce Type: new Abstract: Large language models (LLMs) trained on web-scale corpora can memorize sensitive training data, posing significant privacy risks. Differential privacy (

safetyarxiv-cs-cl
13 May 2026
Safety

How far can bias go? Tracing bias from pretraining data to alignment

DGX agent

arXiv:2411.19240v2 Announce Type: replace Abstract: As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much wo

safetyarxiv-cs-cl
13 May 2026
Model Releases

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

DGX agent

arXiv:2605.11663v1 Announce Type: new Abstract: Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japan

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

DGX agent

arXiv:2505.13770v3 Announce Type: replace-cross Abstract: Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it rema

model-releasesarxiv-cs-cl
13 May 2026
Applications

Instructions shape Production of Language, not Processing

DGX agent

arXiv:2605.11206v1 Announce Type: new Abstract: Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and prod

applicationsarxiv-cs-cl
13 May 2026
Safety

Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation

DGX agent

arXiv:2510.17062v2 Announce Type: replace Abstract: While reasoning-based large language models excel at complex tasks through an internal, structured thinking process, a concerning phenomenon has eme

safetyarxiv-cs-cl
13 May 2026
Safety

Invisible failures in human-AI interactions

DGX agent

arXiv:2603.15423v2 Announce Type: replace Abstract: AI systems fail silently far more often than they fail visibly. In an analysis of 100K human-AI interactions from the WildChat dataset, we find that

safetyarxiv-cs-cl
13 May 2026
Research

Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition

DGX agent

arXiv:2605.12047v1 Announce Type: new Abstract: Is child-directed language (CDL) optimized to support language learning, and which aspects of linguistic development does it facilitate? We investigate

researcharxiv-cs-cl
13 May 2026
Model Releases

KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

DGX agent

arXiv:2605.12471v1 Announce Type: cross Abstract: We introduce KV-Fold, a simple, training-free long-context inference protocol that treats the key-value (KV) cache as the accumulator in a left fold o

model-releasesarxiv-cs-cl
13 May 2026
Research

Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence

DGX agent

arXiv:2605.11348v1 Announce Type: new Abstract: During disasters, extracting causal relations from social media can strengthen situational awareness by identifying factors linked to casualties, physic

researcharxiv-cs-cl
13 May 2026
Model Releases

Latent Causal Void: Explicit Missing-Context Reconstruction for Misinformation Detection

DGX agent

arXiv:2605.12156v1 Announce Type: new Abstract: Automatic misinformation detection performs well when deception is visible in what an article explicitly states. However, some misinformation articles r

model-releasesarxiv-cs-cl
13 May 2026
Safety

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

DGX agent

arXiv:2605.11301v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, c

safetyarxiv-cs-cl
13 May 2026
Safety

Learning Adapter Rank via Symmetry Breaking

DGX agent

arXiv:2506.22809v4 Announce Type: replace-cross Abstract: Low-rank adaptation is effective partly because downstream updates lie in a low-dimensional subspace, but the latent rank coordinates of LoRA

safetyarxiv-cs-cl
13 May 2026
Safety

Learning Agentic Policy from Action Guidance

DGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

safetyarxiv-cs-cl
13 May 2026
Model Releases

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation

DGX agent

arXiv:2605.11739v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute t

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

DGX agent

arXiv:2605.12493v1 Announce Type: new Abstract: Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workfl

model-releasesarxiv-cs-cl
13 May 2026
Safety

MajinBook: An open catalogue of digitally mediated world literature

DGX agent

arXiv:2511.11412v5 Announce Type: replace Abstract: This data paper introduces MajinBook, an open catalogue designed to facilitate the use of shadow libraries-such as Library Genesis and Z-Library-for

safetyarxiv-cs-cl
13 May 2026
Applications

MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification

DGX agent

arXiv:2605.11408v1 Announce Type: cross Abstract: Tabular data forms the backbone of high-stakes decision systems in finance, healthcare, and beyond. Yet industrial tabular datasets are inherently dif

applicationsarxiv-cs-cl
13 May 2026
Applications

Mechanistic Interpretability of ASR models using Sparse Autoencoders

DGX agent

arXiv:2605.12225v1 Announce Type: new Abstract: Understanding the internal machinations of deep Transformer-based NLP models is more crucial than ever as these models see widespread use in various dom

applicationsarxiv-cs-cl
13 May 2026
Model Releases

MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering

DGX agent

arXiv:2605.12361v1 Announce Type: new Abstract: Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain dis

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

MEME: Multi-entity & Evolving Memory Evaluation

DGX agent

arXiv:2605.12477v1 Announce Type: cross Abstract: LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. Whil

model-releasesarxiv-cs-cl
13 May 2026
Safety

Metaphor Is Not All Attention Needs

DGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

safetyarxiv-cs-cl
13 May 2026
Tutorials

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

DGX agent

arXiv:2605.12242v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and

tutorialsarxiv-cs-cl
13 May 2026
Model Releases

Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding

DGX agent

arXiv:2605.12185v1 Announce Type: new Abstract: Large language models accumulate extensive parametric knowledge through pre-training. However, knowledge conflicts occur when outdated or incorrect para

model-releasesarxiv-cs-cl
13 May 2026
Research

MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

DGX agent

arXiv:2605.05945v2 Announce Type: replace-cross Abstract: The recent advancement of Vision Language Action (VLA) models has driven a critical demand for large scale egocentric datasets. However, exist

researcharxiv-cs-cl
13 May 2026
Model Releases

Modality-Inconsistent Continual Learning of Multimodal Large Language Models

DGX agent

arXiv:2412.13050v2 Announce Type: replace-cross Abstract: In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language M

model-releasesarxiv-cs-cl
13 May 2026
Research

Modeling Narrative Structure in Latin Epic Poetry with Automatically Generated Story Grammars

DGX agent

arXiv:2502.12276v2 Announce Type: replace Abstract: Computational methods for analyzing prose and poetry utilize word embeddings and other abstract representations that sometimes obscure context-rich

researcharxiv-cs-cl
13 May 2026
Model Releases

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

DGX agent

arXiv:2605.11836v1 Announce Type: cross Abstract: Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabiliti

model-releasesarxiv-cs-cl
13 May 2026
Research

Much of Geospatial Web Search Is Beyond Traditional GIS

DGX agent

arXiv:2605.11336v1 Announce Type: cross Abstract: Web search queries concern place far more often than existing labelling schemes suggest, yet the landscape of geospatial web search queries - what peo

researcharxiv-cs-cl
13 May 2026
Agents

Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs

DGX agent

arXiv:2605.12460v1 Announce Type: cross Abstract: The continued improvements in language model capability have unlocked their widespread use as drivers of autonomous agents, for example in coding or c

agentsarxiv-cs-cl
13 May 2026
Safety

Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models

DGX agent

arXiv:2605.11959v1 Announce Type: cross Abstract: Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features

safetyarxiv-cs-cl
13 May 2026
Tutorials

Natural Language Processing in the Legal Domain

DGX agent

arXiv:2302.12039v2 Announce Type: replace Abstract: We summarize the current state of the field of NLP & Law with a specific focus on recent technical and substantive developments. To support our anal

tutorialsarxiv-cs-cl
13 May 2026
Model Releases

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation

DGX agent

arXiv:2605.12207v1 Announce Type: cross Abstract: We study the extit{parameter placement problem}: given a fixed budget of k trainable entries within the B matrix of a LoRA adapter (A frozen), does th

model-releasesarxiv-cs-cl
13 May 2026
← Previous
1…9495969798…161
Next →