AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Safety

PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

DGX agent

arXiv:2504.17052v4 Announce Type: replace Abstract: Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspect

safetyarxiv-cs-cl
28 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

DGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

DGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Hard Decision Layer: Evidence for Committed Inference in Transformers

DGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

DGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

DGX agent

arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

DGX agent

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true.

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

DGX agent

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly

model-releasesarxiv-cs-ai
15 Jul 2026
Research

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

DGX agent

arXiv:2607.12297v1 Announce Type: new Abstract: The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applicat

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

DGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

DGX agent

arXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

DGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Validity of LLMs as data annotators: AMALIA on authority

DGX agent

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicl

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Can Reinforcement Learning Efficiently Discover Price Manipulation?

DGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

DGX agent

arXiv:2601.21433v2 Announce Type: replace Abstract: Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framin

safetyarxiv-cs-ai
9 Jul 2026
Model Releases

Predicting LLM Safety Before Release by Simulating Deployment

DGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Enhancement of E-commerce Sponsored Search Relevancy with LLM

DGX agent

arXiv:2607.03886v1 Announce Type: cross Abstract: Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align with the us

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

ICR-RL: Deep Reinforcement Learning via In-Context Regression

DGX agent

arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

NEST: Nascent Encoded Steganographic Thoughts

DGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems

DGX agent

arXiv:2606.19145v2 Announce Type: replace-cross Abstract: Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechan

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges

DGX agent

arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Neural Network-Based Estimation of Time-Dependent Parameters in AR(p) Processes

DGX agent

arXiv:2607.00470v1 Announce Type: cross Abstract: We investigate a forecasting framework based on a simple discrete-time dynamic model with coefficients varying in time. The parameters of the model ar

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

DGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

DGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Training-free Truthfulness Detection via Sparse MLP Value Vectors

DGX agent

arXiv:2509.17932v2 Announce Type: replace Abstract: Large language models (LLMs) are prone to generating factually incorrect content, motivating methods for assessing truthfulness from internal model

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

DGX agent

arXiv:2606.26618v1 Announce Type: new Abstract: Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their traini

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

DGX agent

arXiv:2606.26654v1 Announce Type: new Abstract: Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogu

model-releasesarxiv-cs-cl
26 Jun 2026
Research

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

DGX agent

arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by sh

researcharxiv-cs-cl
25 Jun 2026
Model Releases

Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs

DGX agent

arXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

You Don't Need to Run Every Eval

DGX agent

arXiv:2606.24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

DGX agent

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that prov

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Essential Subspace Merging for Multi-Task Learning

DGX agent

arXiv:2606.19164v2 Announce Type: replace Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

DGX agent

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these d

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

DGX agent

arXiv:2606.22889v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. Howe

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

DGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

DGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

model-releasesarxiv-cs-ai
10 Jun 2026
Applications

A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers

DGX agent

arXiv:2606.07593v1 Announce Type: cross Abstract: The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances o

applicationsarxiv-cs-ai
9 Jun 2026
Model Releases

Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB

DGX agent

arXiv:2511.11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding e decomposes as ilde e + mu, where the

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

DGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

DGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

DGX agent

arXiv:2606.02981v1 Announce Type: new Abstract: Best-of-N inference scaling (drawing N candidate answers from a language model and returning the one a reward model ranks highest) improves accuracy by

researcharxiv-cs-cl
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SenseJudge: Human-Centric Preference-Driven Judgment Framework

DGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

DGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

DGX agent

arXiv:2606.00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Physics-Guided Recurrent State-Space Neural Networks for Multi-Step Prediction

DGX agent

arXiv:2606.02278v1 Announce Type: cross Abstract: State-space models are traditionally based on physical knowledge, but multi-step predictions from these physical models can be poor due to model inacc

researcharxiv-cs-lg
2 Jun 2026
Model Releases

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

DGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

model-releasesarxiv-cs-ai
2 Jun 2026
← Previous
1…197198199200201…1030
Next →