AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

DGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

model-releasesarxiv-cs-cl
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

DGX agent

arXiv:2607.02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpo

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

DGX agent

arXiv:2603.17673v2 Announce Type: replace-cross Abstract: LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to repro

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking

DGX agent

arXiv:2410.15135v5 Announce Type: replace Abstract: With the surge of online misinformation, Large Language Models (LLMs) and Reasoning Large Language Models (RLMs) serving as Automatic Fact-Checking

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

UniVideo: Unified Understanding, Generation, and Editing for Videos

DGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms

DGX agent

arXiv:2603.29466v2 Announce Type: replace-cross Abstract: Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Distributionally Robust Listwise Preference Optimization

DGX agent

arXiv:2607.01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, o

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

DGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Meta-Benchmarks for Financial-Services LLM Evaluation

DGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

DGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

DGX agent

arXiv:2607.01612v1 Announce Type: new Abstract: Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication

DGX agent

arXiv:2607.01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, w

model-releasesarxiv-cs-lg
3 Jul 2026
Safety

Continuous Speculative Decoding for Autoregressive Image Generation

DGX agent

arXiv:2411.11925v3 Announce Type: replace Abstract: Continuous visual autoregressive (AR) models have demonstrated promising performance in image generation, but their inherently sequential nature res

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

DGX agent

arXiv:2607.00341v1 Announce Type: cross Abstract: Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). How

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

DGX agent

arXiv:2512.15702v2 Announce Type: replace Abstract: Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. Wh

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

From 'Strings' to 'Things' for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems

DGX agent

arXiv:2607.00003v1 Announce Type: cross Abstract: Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentra

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data

DGX agent

arXiv:2607.00733v1 Announce Type: cross Abstract: Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of under

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

DGX agent

arXiv:2607.00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Exis

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection

DGX agent

arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

DGX agent

arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps

DGX agent

arXiv:2607.00004v1 Announce Type: cross Abstract: While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the agi

model-releasesarxiv-cs-ai
2 Jul 2026
Research

Amplifying Membership Signal Through Chained Regeneration

DGX agent

arXiv:2606.31991v1 Announce Type: cross Abstract: The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enforcement. C

researcharxiv-cs-ai
1 Jul 2026
Model Releases

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

DGX agent

arXiv:2606.30850v1 Announce Type: new Abstract: Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic unce

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

DGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

model-releasesarxiv-cs-cl
1 Jul 2026
Local Ai

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

DGX agent

arXiv:2606.31844v1 Announce Type: cross Abstract: A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observab

local-aiarxiv-cs-ai
1 Jul 2026
Applications

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

DGX agent

arXiv:2606.31332v1 Announce Type: new Abstract: Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity.

applicationsarxiv-cs-ai
1 Jul 2026
Applications

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

DGX agent

arXiv:2606.31257v1 Announce Type: new Abstract: The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-la

applicationsarxiv-cs-cv
1 Jul 2026
Model Releases

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

DGX agent

arXiv:2606.30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models ha

model-releasesarxiv-cs-lg
1 Jul 2026
Safety

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

DGX agent

arXiv:2509.12046v2 Announce Type: replace-cross Abstract: Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned gen

safetyarxiv-cs-ai
1 Jul 2026
Model Releases

Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias

DGX agent

arXiv:2606.30821v1 Announce Type: new Abstract: Probabilistic downscaling is the task of modeling the conditional distribution of high-resolution fields given coarse inputs, and is a central challenge

model-releasesarxiv-cs-lg
1 Jul 2026
Agents

Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving

DGX agent

arXiv:2606.31160v1 Announce Type: new Abstract: Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step before producing a trajectory. The re

agentsarxiv-cs-cv
1 Jul 2026
Local Ai

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

DGX agent

arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them fro

local-aiarxiv-cs-ai
1 Jul 2026
Research

Stage-Transition Dense Reward Modeling for Reinforcement Learning

DGX agent

arXiv:2606.31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping si

researcharxiv-cs-ai
1 Jul 2026
Model Releases

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

DGX agent

arXiv:2606.31039v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies re

model-releasesarxiv-cs-cl
1 Jul 2026
Research

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

DGX agent

arXiv:2606.31128v1 Announce Type: cross Abstract: Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-lev

researcharxiv-cs-ai
1 Jul 2026
Safety

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

DGX agent

arXiv:2603.16271v3 Announce Type: replace Abstract: Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial d

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

DGX agent

arXiv:2606.29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observati

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

BERTomelo: Your Portuguese Encoder Best Friend

DGX agent

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

DGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

DGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

model-releasesarxiv-cs-cl
30 Jun 2026
Safety

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

DGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

DGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

model-releasesarxiv-cs-cl
30 Jun 2026
Research

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

DGX agent

arXiv:2606.29897v1 Announce Type: cross Abstract: Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are

researcharxiv-cs-ai
30 Jun 2026
Model Releases

Demonstration-Free Robotic Control via LLM Agents

DGX agent

arXiv:2601.20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task

model-releasesarxiv-cs-ai
30 Jun 2026
Research

DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent

DGX agent

arXiv:2606.28417v1 Announce Type: new Abstract: Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To co

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Diversity is the Strength of the AI Crowd

DGX agent

arXiv:2606.29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combine

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

DGX agent

arXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

DGX agent

arXiv:2510.25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits. However,

model-releasesarxiv-cs-ai
30 Jun 2026
← Previous
1…280281282283284…1058
Next →