AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs

DGX agent

arXiv:2606.01879v1 Announce Type: new Abstract: Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their a

model-releasesarxiv-cs-cl
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

FACT: A Simple and Efficient Framework for Active Finetuning

DGX agent

arXiv:2606.02079v1 Announce Type: new Abstract: The main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Feature to Dynamics: Feature-space to Autoregression strategy for Zero-shot Time Series Forecasting

DGX agent

arXiv:2606.01289v1 Announce Type: new Abstract: Zero-shot time series forecasting aims to predict future values for previously unseen series, requiring models to generalize temporal dynamics beyond th

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction

DGX agent

arXiv:2511.12081v2 Announce Type: replace-cross Abstract: Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark

model-releasesarxiv-cs-lg
2 Jun 2026
Research

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

DGX agent

arXiv:2606.00110v1 Announce Type: new Abstract: Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordi

researcharxiv-cs-cv
2 Jun 2026
Safety

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

DGX agent

arXiv:2510.18439v3 Announce Type: replace Abstract: Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

DGX agent

arXiv:2606.01132v1 Announce Type: new Abstract: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchma

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs

DGX agent

arXiv:2606.00642v1 Announce Type: new Abstract: Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular

researcharxiv-cs-ai
2 Jun 2026
Model Releases

How to Correctly Report LLM-as-a-Judge Evaluations

DGX agent

arXiv:2511.21140v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensiti

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Improving IoT Intrusion Detection Through SMOTE-Based Oversampling and Extended Multi-Model Evaluation on Side-Channel Power Data

DGX agent

arXiv:2606.00161v1 Announce Type: cross Abstract: The detection of intrusions in IoT-based networks poses challenges that cannot be overcome using traditional machine learning methods. Perhaps the big

researcharxiv-cs-ai
2 Jun 2026
Safety

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

DGX agent

arXiv:2606.00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models

safetyarxiv-cs-ai
2 Jun 2026
Research

Logit Distillation on Manifolds: Mapping by Learning

DGX agent

arXiv:2606.00771v1 Announce Type: cross Abstract: A simple way to improve the performance of almost any machine learning model is not to train a single but several models with diverse algorithms which

researcharxiv-cs-ai
2 Jun 2026
Model Releases

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

DGX agent

arXiv:2606.02470v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and too

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

PaintBench: Deterministic Evaluation of Precise Visual Editing

DGX agent

arXiv:2606.00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To p

model-releasesarxiv-cs-lg
2 Jun 2026
Local Ai

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

DGX agent

arXiv:2606.00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising la

local-aiarxiv-cs-ai
2 Jun 2026
Model Releases

ProductWebGen: Benchmarking Multimodal Product Webpage Generation

DGX agent

arXiv:2606.01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value f

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression

DGX agent

arXiv:2606.00494v1 Announce Type: new Abstract: Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. Ho

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

DGX agent

arXiv:2601.04946v3 Announce Type: replace-cross Abstract: Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

DGX agent

arXiv:2606.01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling

model-releasesarxiv-cs-lg
2 Jun 2026
Local Ai

Structure Enables Effective Self-Localization of Errors in LLMs

DGX agent

arXiv:2602.02416v2 Announce Type: replace Abstract: Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect rea

local-aiarxiv-cs-ai
2 Jun 2026
Research

Subliminal Learning Is Steering Vector Distillation

DGX agent

arXiv:2606.00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the

researcharxiv-cs-ai
2 Jun 2026
Research

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

DGX agent

arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from s

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights

DGX agent

arXiv:2507.09766v2 Announce Type: replace-cross Abstract: Accurate estimation of Remaining Useful Life (RUL) and State of Health (SoH) is essential for reliable Prognostics and Health Management (PHM)

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

DGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

DGX agent

arXiv:2606.01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, c

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Variational Learning for Insertion-based Generation

DGX agent

arXiv:2606.02133v1 Announce Type: cross Abstract: Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by

researcharxiv-cs-ai
2 Jun 2026
Research

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

DGX agent

arXiv:2606.02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (de

researcharxiv-cs-ai
2 Jun 2026
Model Releases

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

DGX agent

arXiv:2605.30469v1 Announce Type: cross Abstract: 3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

DGX agent

arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for Engl

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Automating Formal Verification with Reinforcement Learning and Recursive Inference

DGX agent

arXiv:2605.30914v1 Announce Type: new Abstract: Automated formal verification remains challenging for large language models because data for proof assistants and verification-aware languages is scarce

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Can LLM Teams Play What? Where? When?

DGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

DGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

DGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

DGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

DGX agent

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appro

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

Forecasting with Hyper-Trees

DGX agent

arXiv:2405.07836v5 Announce Type: replace Abstract: We introduce Hyper-Trees as a novel framework for modeling time series data using gradient boosted trees. Unlike conventional tree-based approaches

safetyarxiv-cs-lg
1 Jun 2026
Research

Generative Models and Statistical Validation

DGX agent

arXiv:2605.30453v1 Announce Type: cross Abstract: Generative machine learning has become an essential tool in theoretical and experimental physics, especially in the context of fast surrogates and den

researcharxiv-cs-lg
1 Jun 2026
Research

Graph Machine Learning in the Era of Large Language Models (LLMs)

DGX agent

arXiv:2404.14928v3 Announce Type: replace-cross Abstract: Graphs play an important role in representing complex relationships in various domains like social networks, knowledge graphs, and molecular d

researcharxiv-cs-ai
1 Jun 2026
Research

Modeling Covariate Transition for Efficient Estimation of Longitudinal Treatment Effects in Randomized Experiments

DGX agent

arXiv:2605.31443v1 Announce Type: cross Abstract: We present a regression-adjustment framework designed for the estimation of longitudinal treatment effects in randomized experiments under static regi

researcharxiv-cs-lg
1 Jun 2026
Research

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

DGX agent

arXiv:2605.31393v1 Announce Type: cross Abstract: Sign language translation (SLT) remains constrained by limited paired sign-video/text corpora and heavy-tailed target vocabularies. We study target-si

researcharxiv-cs-ai
1 Jun 2026
Model Releases

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

DGX agent

arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a si

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

DGX agent

arXiv:2605.30911v1 Announce Type: cross Abstract: Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinat

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

DGX agent

arXiv:2605.30788v1 Announce Type: cross Abstract: We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurat

model-releasesarxiv-cs-ai
1 Jun 2026
Applications

Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

DGX agent

arXiv:2605.29278v1 Announce Type: new Abstract: As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We pr

applicationsarxiv-cs-cl
29 May 2026
Model Releases

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

DGX agent

arXiv:2605.29400v1 Announce Type: new Abstract: We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Acti

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Benchmarking Single-Factor Physical Video-to-Audio Generation

DGX agent

arXiv:2605.30339v1 Announce Type: new Abstract: Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical process

model-releasesarxiv-cs-cv
29 May 2026
Safety

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

DGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

safetyarxiv-cs-ai
29 May 2026
Research

Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation

DGX agent

arXiv:2605.30269v1 Announce Type: new Abstract: Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individ

researcharxiv-cs-cv
29 May 2026
← Previous
1…285286287288289…1058
Next →