AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

Distilling Bayesian Belief States into Language Models for Auditable Negotiation

DGX agent

arXiv:2605.04507v1 Announce Type: new Abstract: Negotiation agents must infer what their counterpart values, update those beliefs over dialogue turns, and choose actions under uncertainty. End-to-end

safetyarxiv-cs-cl
7 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation

DGX agent

arXiv:2605.04458v1 Announce Type: new Abstract: Evaluation of long-form, citation-backed reports has lately received significant attention due to the wide-scale adoption of retrieval-augmented generat

researcharxiv-cs-cl
7 May 2026
Safety

Elicitation Matters: How Prompts and Query Protocols Shape LLM Surrogates under Sparse Observations

DGX agent

arXiv:2605.04764v1 Announce Type: new Abstract: Large language models are increasingly used as surrogate models for low-data optimization, but their optimizer-facing prediction and its uncertainty rem

safetyarxiv-cs-cl
7 May 2026
Model Releases

Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation

DGX agent

arXiv:2505.18244v3 Announce Type: replace Abstract: Why do language models from different architecture families respond so differently to the same perturbation? We argue that the answer is not scale,

model-releasesarxiv-cs-cl
7 May 2026
Safety

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content

DGX agent

arXiv:2605.04085v1 Announce Type: cross Abstract: Objectives: Large language models (LLMs) are increasingly used for clinical text summarization, yet structured methods to assess associated patient sa

safetyarxiv-cs-cl
7 May 2026
Safety

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL

DGX agent

arXiv:2605.04719v1 Announce Type: new Abstract: Tool-integrated Text-to-SQL parsing has emerged as a promising paradigm, framing SQL generation as a sequential decision-making process interleaved with

safetyarxiv-cs-cl
7 May 2026
Research

FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation

DGX agent

arXiv:2605.04651v1 Announce Type: cross Abstract: Adapting pretrained models typically involves a trade-off between the high training costs of backpropagation and the heavy inference overhead of memor

researcharxiv-cs-cl
7 May 2026
Research

FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals

DGX agent

arXiv:2605.04157v1 Announce Type: new Abstract: SemEval-2026 Task 13 investigates machine-generated code detection across multiple programming languages and application scenarios, asking participating

researcharxiv-cs-cl
7 May 2026
Model Releases

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2

DGX agent

arXiv:2512.22671v2 Announce Type: replace Abstract: Structured width pruning of GLU-MLP layers, guided by the Maximum Absolute Weight (MAW) criterion, reveals a systematic dichotomy in how reducing th

model-releasesarxiv-cs-cl
7 May 2026
Tutorials

France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions

DGX agent

arXiv:2602.23547v2 Announce Type: replace Abstract: Sentences like 'She will go to France or Spain, or perhaps to Germany or France.' appear formally redundant, yet become acceptable in contexts such

tutorialsarxiv-cs-cl
7 May 2026
Model Releases

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

DGX agent

arXiv:2605.04065v1 Announce Type: new Abstract: Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, exi

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents

DGX agent

arXiv:2604.01496v2 Announce Type: replace-cross Abstract: We introduce SWE-ZERO to SWE-HERO, a two-stage SFT recipe that achieves state-of-the-art results on SWE-bench by distilling open-weight fronti

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

DGX agent

arXiv:2605.04135v1 Announce Type: cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequenti

model-releasesarxiv-cs-cl
7 May 2026
Agents

GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking

DGX agent

arXiv:2605.04449v1 Announce Type: new Abstract: Dialogue State Tracking (DST) requires precise extraction of structured information from multi-domain conversations, a task where Large Language Models

agentsarxiv-cs-cl
7 May 2026
Model Releases

Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models

DGX agent

arXiv:2605.04638v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is an important technique for ensuring the trustworthiness of LLMs, given their tendency to hallucinate. Existing state-

model-releasesarxiv-cs-cl
7 May 2026
Safety

Graph-Augmented LLMs for Swiss MP Ideology Prediction

DGX agent

arXiv:2605.04643v1 Announce Type: new Abstract: Approximating the ideological position of Members of Parliament (MPs) is a fundamental task in political science, helping researchers understand legisla

safetyarxiv-cs-cl
7 May 2026
Research

Gyan: An Explainable Neuro-Symbolic Language Model

DGX agent

arXiv:2605.04759v1 Announce Type: new Abstract: Transformer based pre-trained large language models have become ubiquitous. There is increasing evidence to suggest that even with large scale pre-train

researcharxiv-cs-cl
7 May 2026
Research

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

DGX agent

arXiv:2605.04500v1 Announce Type: new Abstract: Low-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models. A great deal of cross-lingu

researcharxiv-cs-cl
7 May 2026
Local Ai

HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search

DGX agent

arXiv:2605.04103v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) has emerged as a powerful framework for automatically discovering neural architectures that balance accuracy and effi

local-aiarxiv-cs-cl
7 May 2026
Agents

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

DGX agent

arXiv:2605.04304v1 Announce Type: cross Abstract: Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While ex

agentsarxiv-cs-cl
7 May 2026
Research

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research

DGX agent

arXiv:2507.15736v2 Announce Type: replace Abstract: Innovation is a key driving force of human civilization. As the body of knowledge has grown considerably, bridging knowledge across different discip

researcharxiv-cs-cl
7 May 2026
Research

Implicit Representations of Grammaticality in Language Models

DGX agent

arXiv:2605.05197v1 Announce Type: new Abstract: Grammaticality and likelihood are distinct notions in human language. Pretrained language models (LMs), which are probabilistic models of language fitte

researcharxiv-cs-cl
7 May 2026
Model Releases

Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks

DGX agent

arXiv:2605.04217v1 Announce Type: cross Abstract: Relative positional encodings determine which functions of query-key lag can enter the primitive attention logit. RoPE supplies a rotary phase, while

model-releasesarxiv-cs-cl
7 May 2026
Safety

LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey

DGX agent

arXiv:2505.00753v5 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-bas

safetyarxiv-cs-cl
7 May 2026
Model Releases

Making Knowledge Accessible: Divergent Readability-Accuracy Strategies of Mistral and QWen in Biomedical Text Simplification

DGX agent

arXiv:2511.05080v4 Announce Type: replace Abstract: The growing public demand for accessible biomedical information calls for scalable text simplification. While large language models (LLMs) offer sol

model-releasesarxiv-cs-cl
7 May 2026
Agents

Material Database Agent: A Multimodal Agentic Framework for Scientific Literature Mining

DGX agent

arXiv:2605.04278v1 Announce Type: new Abstract: Materials science workflows rely on structured and unstructured data from the vast body of available scientific literature. However, most of the experim

agentsarxiv-cs-cl
7 May 2026
Research

Measuring Psychological States Through Semantic Projection: A Theory-Driven Approach to Language-Based Assessment

DGX agent

arXiv:2605.04873v1 Announce Type: new Abstract: Recent advances in natural language processing have enabled increasingly accurate estimation of psychological traits from language. However, most existi

researcharxiv-cs-cl
7 May 2026
Safety

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs

DGX agent

arXiv:2605.04180v1 Announce Type: new Abstract: Large Language Models exhibit strong reasoning and semantic understanding capabilities but often hallucinate in domains that require expert knowledge, a

safetyarxiv-cs-cl
7 May 2026
Safety

Misaligned by Reward: Socially Undesirable Preferences in LLMs

DGX agent

arXiv:2605.05003v1 Announce Type: new Abstract: Reward models are a key component of large language model alignment, serving as proxies for human preferences during training. However, existing evaluat

safetyarxiv-cs-cl
7 May 2026
Model Releases

MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge

DGX agent

arXiv:2605.05175v1 Announce Type: cross Abstract: Background: Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions, where top proprietary models already score highly, limit

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise

DGX agent

arXiv:2605.04313v1 Announce Type: new Abstract: Causal reasoning in natural language requires identifying relevant variables, understanding their interactions, and reasoning about effects and interven

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Not All That Is Fluent Is Factual: Investigating Hallucinations of Large Language Models in Academic Writing

DGX agent

arXiv:2605.04171v1 Announce Type: new Abstract: Large Language models (LLMs) show extraordinary abilities, but they are still prone to hallucinations, especially when we use them for generating Academ

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

DGX agent

arXiv:2605.04208v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive multilingual capabilities for well-resourced languages, yet their performance on low-resource

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Open-Source Image Editing Models Are Zero-Shot Vision Learners

DGX agent

arXiv:2605.04566v1 Announce Type: cross Abstract: Recent studies have shown that large generative models can solve vision tasks they were not explicitly trained for. However, existing evidence relies

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs

DGX agent

arXiv:2605.04665v1 Announce Type: new Abstract: When the substantive content of a request is rewritten, do large language models still answer in the format the original task asked for? We find that th

model-releasesarxiv-cs-cl
7 May 2026
Research

Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

DGX agent

arXiv:2605.04127v1 Announce Type: cross Abstract: Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern

researcharxiv-cs-cl
7 May 2026
Safety

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection

DGX agent

arXiv:2601.09195v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) is a fundamental post-training strategy to align Large Language Models (LLMs) with human intent. However, traditional S

safetyarxiv-cs-cl
7 May 2026
Agents

ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation

DGX agent

arXiv:2510.25224v3 Announce Type: replace Abstract: While Large Language Models (LLMs) are increasingly used in agentic frameworks to assist individual users, there is a growing need for agents that c

agentsarxiv-cs-cl
7 May 2026
Model Releases

PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation

DGX agent

arXiv:2605.05159v1 Announce Type: new Abstract: We present our system for SemEval-2026 Task 9: Multilingual Polarization Detection, a binary classification task spanning 22 languages. Our approach fin

model-releasesarxiv-cs-cl
7 May 2026
Safety

Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations

DGX agent

arXiv:2505.18466v2 Announce Type: replace Abstract: Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multiling

safetyarxiv-cs-cl
7 May 2026
Research

RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation

DGX agent

arXiv:2605.04523v1 Announce Type: new Abstract: We present our winning system for Task~B (generation with reference passages) in SemEval-2026 Task~8: MTRAGEval. Our method is a heterogeneous ensemble

researcharxiv-cs-cl
7 May 2026
Safety

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments

DGX agent

arXiv:2508.04204v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content g

safetyarxiv-cs-cl
7 May 2026
Safety

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization

DGX agent

arXiv:2605.04920v1 Announce Type: cross Abstract: Compositional generalization refers to correctly interpret novel combinations of known primitives, which remains a major challenge. Existing approache

safetyarxiv-cs-cl
7 May 2026
Research

RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction

DGX agent

arXiv:2605.04075v1 Announce Type: cross Abstract: Multimodal Large Language Models face severe challenges in computational efficiency and memory consumption due to the substantial expansion of the vis

researcharxiv-cs-cl
7 May 2026
Local Ai

Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training

DGX agent

arXiv:2605.04913v1 Announce Type: new Abstract: LLM post-training typically propagates task gradients through the full depth of the model. Although this end-to-end structure is simple and general, it

local-aiarxiv-cs-cl
7 May 2026
Model Releases

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

DGX agent

arXiv:2605.04539v1 Announce Type: new Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference si

model-releasesarxiv-cs-cl
7 May 2026
Research

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States

DGX agent

arXiv:2605.04496v1 Announce Type: new Abstract: Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can

researcharxiv-cs-cl
7 May 2026
Model Releases

Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

DGX agent

arXiv:2605.04893v1 Announce Type: cross Abstract: Large language models hallucinate in predictable ways: attention routing fails by over-concentrating on a narrow set of positions, or by spreading so

model-releasesarxiv-cs-cl
7 May 2026
← Previous
1…104105106107108…161
Next →