AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

DGX agent

arXiv:2605.30637v1 Announce Type: new Abstract: Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future he

model-releasesarxiv-cs-ai
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

DGX agent

arXiv:2605.30690v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. Howeve

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

DGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

DGX agent

arXiv:2605.31170v1 Announce Type: cross Abstract: Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

End-to-End Compression for Tabular Foundation Models

DGX agent

arXiv:2602.05649v2 Announce Type: replace Abstract: The long-standing dominance of gradient-boosted decision trees for tabular data has recently been challenged by in-context learning tabular foundati

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

DGX agent

arXiv:2605.31251v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to th

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Error Amplification Limits ANN-to-SNN Conversion in Continuous Control

DGX agent

arXiv:2601.21778v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) can achieve competitive performance by converting already existing well-trained Artificial Neural Networks (ANN

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

DGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy

DGX agent

arXiv:2605.31506v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is the current industry standard for grounding AI in real-world facts. Traditional retrieval methods rely on keyw

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

EvoDefense: Co-Evolving Black-Box Defense with Large Language Models

DGX agent

arXiv:2605.31140v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Expand Neurons, Not Parameters

DGX agent

arXiv:2510.04500v2 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves perform

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Expert Merging in Sparse Mixture of Experts with Nash Bargaining

DGX agent

arXiv:2510.16138v2 Announce Type: replace Abstract: Existing expert merging strategies for Sparse Mixture of Experts (SMoE) typically rely on input-dependent or input-independent averaging of expert p

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents

DGX agent

arXiv:2605.30712v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong capabilities in reasoning, tool use, and multi-step interaction, but they often solve tasks from scr

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Exploring Autonomous Agentic Data Engineering for Model Specialization

DGX agent

arXiv:2605.30407v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without hig

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

DGX agent

arXiv:2605.30947v1 Announce Type: new Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantit

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

DGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

DGX agent

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appro

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

DGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

DGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Float8@2bits: Entropy Coding Enables Data-Free Model Compression

DGX agent

arXiv:2601.22787v2 Announce Type: replace Abstract: Post-training compression is currently divided into two contrasting regimes. On the one hand, fast, data-free, and model-agnostic methods (e.g., NF4

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

DGX agent

arXiv:2605.31145v1 Announce Type: cross Abstract: In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly wi

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

ForecastCompass: Guiding Agentic Forecasting with Adaptive Factor Memory

DGX agent

arXiv:2605.30858v1 Announce Type: new Abstract: Agentic forecasting is important for decision-making in dynamic environments, but it remains challenging because agents must reason from incomplete, tim

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors

DGX agent

arXiv:2605.31042v1 Announce Type: cross Abstract: LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and wr

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents

DGX agent

arXiv:2603.18382v2 Announce Type: replace Abstract: Anonymization is often assumed to protect privacy once explicit identifiers are removed, because re-identification has historically required special

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring

DGX agent

arXiv:2605.31400v1 Announce Type: new Abstract: Real-world image deblurring demands both high-fidelity restoration and computational efficiency, a balance existing methods often struggle to achieve. I

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Gait2Hip-60: A Unified Deep Learning Benchmark for Predicting Hip Muscle Forces and Joint Moments from Multi-Cadence Gait Kinematics

DGX agent

arXiv:2605.30374v1 Announce Type: new Abstract: Estimating hip muscle forces and joint moments during gait typically relies on musculoskeletal simulation, which is informative but time-consuming and d

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing

DGX agent

arXiv:2509.14221v3 Announce Type: replace-cross Abstract: Generative Engine Marketing (GEM) is an emerging ecosystem for monetizing generative engines, such as LLM-based chatbots, by seamlessly integr

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

DGX agent

arXiv:2605.30529v1 Announce Type: cross Abstract: Sentence-embedding models for semantic search are overwhelmingly developed and evaluated on English corpora. When applied to clinical retrieval in oth

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Generating Graph-like Rules for Knowledge Graph Reasoning via Diffusion Models

DGX agent

arXiv:2605.30747v1 Announce Type: new Abstract: Logical rules constitute a cornerstone of knowledge graph (KG) reasoning, valued for their interpretability and ability to model relational patterns. Ho

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation

DGX agent

arXiv:2605.30984v1 Announce Type: cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output d

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go

DGX agent

arXiv:2511.10868v2 Announce Type: replace Abstract: Training data imbalance poses a major challenge for code LLMs. Most available data heavily over represents raw opensource code while underrepresenti

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Goldfish: Monolingual Language Models for 350 Languages

DGX agent

arXiv:2408.10441v3 Announce Type: replace Abstract: For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Graph-Conditioned Mixture of Graph Neural Network Experts for Traffic Forecasting

DGX agent

arXiv:2605.30486v1 Announce Type: cross Abstract: Spatio-temporal forecasting on sensor graphs is commonly tackled with a single backbone architecture applied uniformly across all nodes, although grap

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

DGX agent

arXiv:2605.31031v1 Announce Type: new Abstract: Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce Gr

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

GUI-C^2: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

DGX agent

arXiv:2605.30884v1 Announce Type: new Abstract: Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

DGX agent

arXiv:2605.30621v1 Announce Type: new Abstract: LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Hedging on the Frontier: Learning New Tasks with Few Samples

DGX agent

arXiv:2605.30997v1 Announce Type: cross Abstract: When a learner faces a new task with few samples, it must leverage any available side information. In practice, this often comes in the form of model

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

DGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

DGX agent

arXiv:2602.09309v2 Announce Type: replace-cross Abstract: Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this th

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

DGX agent

arXiv:2510.15614v3 Announce Type: replace Abstract: Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effe

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Identifiable Equivariant Networks are Layerwise Equivariant

DGX agent

arXiv:2601.21645v2 Announce Type: replace Abstract: We investigate the relation between end-to-end equivariance and layerwise equivariance in deep neural networks. We prove the following: For a networ

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law

DGX agent

arXiv:2605.30589v1 Announce Type: cross Abstract: U.S. immigration law spans thousands of pages of official policy, federal regulations, and procedural guidance that change frequently and carry high s

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback

DGX agent

arXiv:2605.30478v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) trains language models using programmatically checkable signals such as unit-test outcomes, enab

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data

DGX agent

arXiv:2605.31324v1 Announce Type: cross Abstract: Estimating the generalization gap and developing optimization methods that improve generalization are crucial for deep learning models, for both theor

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

DGX agent

arXiv:2605.30917v1 Announce Type: cross Abstract: As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing atte

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Inversion-Free Natural Gradient Descent on Riemannian Manifolds

DGX agent

arXiv:2604.02969v2 Announce Type: replace-cross Abstract: The natural gradient method is a central tool for statistical optimisation, but its broader application is hindered by the assumption of a Euc

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware

DGX agent

arXiv:2603.08721v2 Announce Type: replace-cross Abstract: New AI accelerators with novel instruction set architectures (ISAs) often require developers to manually craft low-level kernels, a time-consu

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation

DGX agent

arXiv:2605.31478v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate power-system analysis, but many utilities and energy-research labs require on-premise s

model-releasesarxiv-cs-cl
1 Jun 2026
← Previous
1…175176177178179…361
Next →