AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
13 May 2026

Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines

Model ReleasesDGX agent

arXiv:2601.03627v3 Announce Type: replace Abstract: We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs

Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs

ResearchDGX agent

arXiv:2505.02072v2 Announce Type: replace Abstract: Language modeling has shifted in recent years from a distribution over strings to prediction models with textual inputs and outputs for general-purp

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.11403v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Opti

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

Model ReleasesDGX agent

arXiv:2504.14707v3 Announce Type: replace Abstract: We introduce FLAME (FLemish Accounts of Momentary Experiences), a new corpus of nearly 25,000 daily personal narratives in Belgian-Dutch (Flemish),

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

Model ReleasesDGX agent

arXiv:2605.11416v1 Announce Type: new Abstract: Selective layer-wise updates are essential for low-cost continued pre-training of Large Language Models (LLMs), yet determining which layers to freeze o

From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction

ApplicationsDGX agent

arXiv:2605.11774v1 Announce Type: new Abstract: By processing electronic health records (EHRs) as natural language sequences, large language models (LLMs) have shown potential in clinical prediction t

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

SafetyDGX agent

arXiv:2605.11853v1 Announce Type: cross Abstract: Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that pr

Geometric Factual Recall in Transformers

Model ReleasesDGX agent

arXiv:2605.12426v1 Announce Type: new Abstract: How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of emb

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender

Model ReleasesDGX agent

arXiv:2605.12299v1 Announce Type: new Abstract: Recent works have analyzed the impact of individual components of neural networks on gendered predictions, often with a focus on mitigating gender bias.

Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction

Model ReleasesDGX agent

arXiv:2305.12678v3 Announce Type: replace Abstract: Multimodal Review Helpfulness Prediction (MRHP) aims to rank product reviews based on predicted helpfulness scores and has been widely applied in e-

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

Model ReleasesDGX agent

arXiv:2605.06152v2 Announce Type: replace-cross Abstract: Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the 'Slingshot Mechanism.' Ex

GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs

Model ReleasesDGX agent

arXiv:2503.06139v2 Announce Type: replace Abstract: Pairwise LLM-as-a-judge evaluation asks the judge to identify the better of two candidate answers. We study a one-line modification that asks for th

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

Model ReleasesDGX agent

arXiv:2601.20255v2 Announce Type: replace-cross Abstract: SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabil

HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model

Model ReleasesDGX agent

arXiv:2605.11255v1 Announce Type: new Abstract: We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Traini

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation

ApplicationsDGX agent

arXiv:2605.11651v1 Announce Type: cross Abstract: Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the fi

How Does Differential Privacy Affect Social Bias in LLMs? A Systematic Evaluation

SafetyDGX agent

arXiv:2605.11195v1 Announce Type: new Abstract: Large language models (LLMs) trained on web-scale corpora can memorize sensitive training data, posing significant privacy risks. Differential privacy (

How far can bias go? Tracing bias from pretraining data to alignment

SafetyDGX agent

arXiv:2411.19240v2 Announce Type: replace Abstract: As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much wo

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

Model ReleasesDGX agent

arXiv:2605.11663v1 Announce Type: new Abstract: Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japan

Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

Model ReleasesDGX agent

arXiv:2505.13770v3 Announce Type: replace-cross Abstract: Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it rema

Instructions shape Production of Language, not Processing

ApplicationsDGX agent

arXiv:2605.11206v1 Announce Type: new Abstract: Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and prod

Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation

SafetyDGX agent

arXiv:2510.17062v2 Announce Type: replace Abstract: While reasoning-based large language models excel at complex tasks through an internal, structured thinking process, a concerning phenomenon has eme

Invisible failures in human-AI interactions

SafetyDGX agent

arXiv:2603.15423v2 Announce Type: replace Abstract: AI systems fail silently far more often than they fail visibly. In an analysis of 100K human-AI interactions from the WildChat dataset, we find that

Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition

ResearchDGX agent

arXiv:2605.12047v1 Announce Type: new Abstract: Is child-directed language (CDL) optimized to support language learning, and which aspects of linguistic development does it facilitate? We investigate

KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

Model ReleasesDGX agent

arXiv:2605.12471v1 Announce Type: cross Abstract: We introduce KV-Fold, a simple, training-free long-context inference protocol that treats the key-value (KV) cache as the accumulator in a left fold o

Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence

ResearchDGX agent

arXiv:2605.11348v1 Announce Type: new Abstract: During disasters, extracting causal relations from social media can strengthen situational awareness by identifying factors linked to casualties, physic

Latent Causal Void: Explicit Missing-Context Reconstruction for Misinformation Detection

Model ReleasesDGX agent

arXiv:2605.12156v1 Announce Type: new Abstract: Automatic misinformation detection performs well when deception is visible in what an article explicitly states. However, some misinformation articles r

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

SafetyDGX agent

arXiv:2605.11301v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, c

Learning Adapter Rank via Symmetry Breaking

SafetyDGX agent

arXiv:2506.22809v4 Announce Type: replace-cross Abstract: Low-rank adaptation is effective partly because downstream updates lie in a low-dimensional subspace, but the latent rank coordinates of LoRA

Learning Agentic Policy from Action Guidance

SafetyDGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation

Model ReleasesDGX agent

arXiv:2605.11739v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute t

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Model ReleasesDGX agent

arXiv:2605.12493v1 Announce Type: new Abstract: Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workfl

MajinBook: An open catalogue of digitally mediated world literature

SafetyDGX agent

arXiv:2511.11412v5 Announce Type: replace Abstract: This data paper introduces MajinBook, an open catalogue designed to facilitate the use of shadow libraries-such as Library Genesis and Z-Library-for

MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification

ApplicationsDGX agent

arXiv:2605.11408v1 Announce Type: cross Abstract: Tabular data forms the backbone of high-stakes decision systems in finance, healthcare, and beyond. Yet industrial tabular datasets are inherently dif

Mechanistic Interpretability of ASR models using Sparse Autoencoders

ApplicationsDGX agent

arXiv:2605.12225v1 Announce Type: new Abstract: Understanding the internal machinations of deep Transformer-based NLP models is more crucial than ever as these models see widespread use in various dom

MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering

Model ReleasesDGX agent

arXiv:2605.12361v1 Announce Type: new Abstract: Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain dis

MEME: Multi-entity & Evolving Memory Evaluation

Model ReleasesDGX agent

arXiv:2605.12477v1 Announce Type: cross Abstract: LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. Whil

Metaphor Is Not All Attention Needs

SafetyDGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

TutorialsDGX agent

arXiv:2605.12242v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and

Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding

Model ReleasesDGX agent

arXiv:2605.12185v1 Announce Type: new Abstract: Large language models accumulate extensive parametric knowledge through pre-training. However, knowledge conflicts occur when outdated or incorrect para

MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

ResearchDGX agent

arXiv:2605.05945v2 Announce Type: replace-cross Abstract: The recent advancement of Vision Language Action (VLA) models has driven a critical demand for large scale egocentric datasets. However, exist

Modality-Inconsistent Continual Learning of Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2412.13050v2 Announce Type: replace-cross Abstract: In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language M

Modeling Narrative Structure in Latin Epic Poetry with Automatically Generated Story Grammars

ResearchDGX agent

arXiv:2502.12276v2 Announce Type: replace Abstract: Computational methods for analyzing prose and poetry utilize word embeddings and other abstract representations that sometimes obscure context-rich

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

Model ReleasesDGX agent

arXiv:2605.11836v1 Announce Type: cross Abstract: Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabiliti

Much of Geospatial Web Search Is Beyond Traditional GIS

ResearchDGX agent

arXiv:2605.11336v1 Announce Type: cross Abstract: Web search queries concern place far more often than existing labelling schemes suggest, yet the landscape of geospatial web search queries - what peo

Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs

AgentsDGX agent

arXiv:2605.12460v1 Announce Type: cross Abstract: The continued improvements in language model capability have unlocked their widespread use as drivers of autonomous agents, for example in coding or c

Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models

SafetyDGX agent

arXiv:2605.11959v1 Announce Type: cross Abstract: Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features

Natural Language Processing in the Legal Domain

TutorialsDGX agent

arXiv:2302.12039v2 Announce Type: replace Abstract: We summarize the current state of the field of NLP & Law with a specific focus on recent technical and substantive developments. To support our anal

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.12207v1 Announce Type: cross Abstract: We study the extit{parameter placement problem}: given a fixed budget of k trainable entries within the B matrix of a LoRA adapter (A frozen), does th

Not Worth Mentioning? A Pilot Study on Salient Proposition Annotation

ResearchDGX agent

arXiv:2603.27358v2 Announce Type: replace Abstract: Despite a long tradition of work on extractive summarization, which by nature aims to recover the most important propositions in a text, little work

OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models

ApplicationsDGX agent

arXiv:2605.11629v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct de

On Predicting the Post-training Potential of Pre-trained LLMs

ResearchDGX agent

arXiv:2605.11978v1 Announce Type: new Abstract: The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. Howev

On Problems of Implicit Context Compression for Software Engineering Agents

AgentsDGX agent

arXiv:2605.11051v1 Announce Type: cross Abstract: LLM-based Software Engineering agents face a critical bottleneck: context length limitations cause failures on complex, long-horizon tasks. One promis

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

SafetyDGX agent

arXiv:2605.05630v2 Announce Type: replace Abstract: Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objec

ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

ResearchDGX agent

arXiv:2605.12419v1 Announce Type: new Abstract: Despite the rapid advancements in large language model (LLM) development, fine-tuning them for specific tasks often results in the catastrophic forgetti

ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models

SafetyDGX agent

arXiv:2605.12446v1 Announce Type: cross Abstract: Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential f

Output Composability of QLoRA PEFT Modules for Plug-and-Play Attribute-Controlled Text Generation

Model ReleasesDGX agent

arXiv:2605.12345v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) techniques offer task-specific fine-tuning at a fraction of the cost of full fine-tuning, but require separate fi

Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering

Model ReleasesDGX agent

arXiv:2605.12313v1 Announce Type: new Abstract: Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple s

Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling

AgentsDGX agent

arXiv:2605.12411v1 Announce Type: cross Abstract: AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant ne

Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

ResearchDGX agent

arXiv:2605.12422v1 Announce Type: new Abstract: Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to suc

Predicting Psychological Well-Being from Spontaneous Speech using LLMs

Model ReleasesDGX agent

arXiv:2605.11303v1 Announce Type: new Abstract: We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. U

← Previous
1…7576777879…129
Next →