AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Entity-Collision: A Stratified Protocol for Attributing Retrieval Lift in Agent Memory

DGX agent

arXiv:2605.29630v1 Announce Type: cross Abstract: End-to-end agent-memory benchmarks report a single hit@k per retriever, confounding lexical leakage (uncontrolled query/gold/distractor entity overlap

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ESPO: Early-Stopping Proximal Policy Optimization

DGX agent

arXiv:2605.29860v1 Announce Type: cross Abstract: When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep ge

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

DGX agent

arXiv:2605.29637v1 Announce Type: new Abstract: Large language models recall knowledge reliably in English but often fail on the same query posed in a lower-resourced language -- a crosslingual consis

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach

DGX agent

arXiv:2511.19316v2 Announce Type: replace-cross Abstract: Recent fine-tuning techniques for diffusion models enable them to reproduce specific image sets, such as particular faces or artistic styles,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation

DGX agent

arXiv:2605.29977v1 Announce Type: new Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by ex

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

DGX agent

arXiv:2605.29874v1 Announce Type: cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibriu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Evolutionary Rule Extraction from Corporate Default Prediction Models

DGX agent

arXiv:2605.29478v1 Announce Type: cross Abstract: Small and medium-sized enterprises (SMEs) represent the majority of firms in most economies and often face financial constraints and higher vulnerabil

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ExCAM: Explainable Cultural Awareness Metrics

DGX agent

arXiv:2605.29897v1 Announce Type: new Abstract: Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Fairness-Aware Federated Learning with Trajectory Shapley Value

DGX agent

arXiv:2605.30336v1 Announce Type: new Abstract: Federated learning is an emerging distributed paradigm that addresses the challenges posed by heterogeneous, privacy-sensitive data. It enables multiple

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging

DGX agent

arXiv:2605.29827v1 Announce Type: new Abstract: Medical image analysis models can exhibit performance disparities across patient subgroups, threatening clinical safety and fairness. Existing methods t

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models

DGX agent

arXiv:2511.11505v3 Announce Type: replace Abstract: Blocking communication presents a major hurdle in running MoEs efficiently in distributed settings. To address this, we present FarSkip-Collective w

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models

DGX agent

arXiv:2605.28896v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has emerged as a widely adopted approach for adapting large language models, yet the internal representational changes induce

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

FedQHD: Closed-Form Function-Space Federated Reinforcement Learning

DGX agent

arXiv:2605.29002v1 Announce Type: new Abstract: Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions

DGX agent

arXiv:2605.29427v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulator

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification

DGX agent

arXiv:2605.29586v1 Announce Type: new Abstract: We introduce FinVerBench, a benchmark and validity study for financial statement verification: determining whether a set of corporate financial statemen

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

DGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

FoRA: Fisher-orthogonal Rank Adaptation for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2605.29317v1 Announce Type: new Abstract: Parameter-efficient fine-tuning(PEFT) has largely focused on LoRA and its accuracy-oriented variants, leaving the original goal of reducing trainable pa

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

DGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

FPLIER: Federated Pathway-Level Information Extractor

DGX agent

arXiv:2605.29587v1 Announce Type: cross Abstract: In transcriptomics, gene-set-aware factorization methods such as the Pathway Level Information Extractor (PLIER) are most effective when trained on la

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

From Sublinear to Linear: Local Convergence in Finite-Width Networks via Locally Polyak-Lojasiewicz Regions

DGX agent

arXiv:2507.21429v3 Announce Type: replace-cross Abstract: We study local linear convergence of gradient descent for finite-width feedforward networks under the squared empirical loss. Prior work shows

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

From XXLTraffic to EvoXXLTraffic: Scaling Traffic Forecasting to Sensor-Evolving Networks

DGX agent

arXiv:2605.29768v1 Announce Type: new Abstract: Existing traffic forecasting benchmarks assume a fixed sensor set, but real road-sensor networks grow continuously as the road network changes year by y

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

DGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver

DGX agent

arXiv:2605.30045v1 Announce Type: new Abstract: Video object removal frequently struggles to simultaneously eliminate target objects and their associated physical effects (e.g., smoke, reflections, li

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

DGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

DGX agent

arXiv:2605.30022v1 Announce Type: cross Abstract: Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and st

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

DGX agent

arXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

DGX agent

arXiv:2605.28848v1 Announce Type: cross Abstract: Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all ch

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GPIC: A Giant Permissive Image Corpus for Visual Generation

DGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Gradient Perturbation: Learning to Perturb Gradients for Adaptive Training

DGX agent

arXiv:2605.29494v1 Announce Type: new Abstract: Deep neural network training involves both forward propagation (from features through logits to loss) and backward propagation (from loss through gradie

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Gram: Assessing sabotage propensities via automated alignment auditing

DGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

DGX agent

arXiv:2605.29668v1 Announce Type: new Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GroundAct: Can LLM Agents Ground Actions in Environmental States?

DGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

DGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

DGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

DGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

DGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

DGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

DGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

DGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

DGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

DGX agent

arXiv:2605.30179v1 Announce Type: cross Abstract: Parameter-efficient adaptation has made LLMs practical for domain prediction, but standard LoRA still relies on a static low-rank update and does not

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Improving Adversarial Robustness of Attribution via Implicit Regularization

DGX agent

arXiv:2605.29983v1 Announce Type: cross Abstract: The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typicall

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Improving Full Waveform Inversion in Large Model Era

DGX agent

arXiv:2603.00377v2 Announce Type: replace Abstract: Full Waveform Inversion (FWI) is a highly nonlinear and ill-posed problem that aims to recover subsurface velocity maps from surface-recorded seismi

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

DGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Inferring the Size of Large Language Models From Popular Text Memorization

DGX agent

arXiv:2605.29223v1 Announce Type: new Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary refere

model-releasesarxiv-cs-lg
29 May 2026
← Previous
1…180181182183184…361
Next →