AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,805 results
Model Releases

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

DGX agent

arXiv:2602.18008v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on s

model-releasesarxiv-cs-ai
2 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

DGX agent

arXiv:2606.00602v1 Announce Type: new Abstract: Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak,

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

ASE-26: a curriculum for agentic software engineering as a discipline

DGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Assessment of Generative Named Entity Recognition in the Era of Large Language Models

DGX agent

arXiv:2601.17898v2 Announce Type: replace Abstract: Named entity recognition (NER) is evolving from a sequence labeling task into a generative paradigm with the rise of large language models (LLMs). W

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

DGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

DGX agent

arXiv:2606.02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

DGX agent

arXiv:2606.02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. W

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

AutoEval Done Right: Using Synthetic Data for Model Evaluation

DGX agent

arXiv:2403.07008v3 Announce Type: replace-cross Abstract: The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data c

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

DGX agent

arXiv:2606.01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form c

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Autopilot-Preserving Residual Q-Learning with HJB-Inspired Finite-Action Risk Filtering for Fixed-Wing UAV Command Supervision

DGX agent

arXiv:2606.01397v1 Announce Type: cross Abstract: A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can d

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning

DGX agent

arXiv:2606.00671v1 Announce Type: new Abstract: We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model function

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

DGX agent

arXiv:2606.02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approac

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Bayesian Inference of Nonlinear Malaria Dynamics in Ghana via an Ensemble Markov Chain Monte Carlo Sampler

DGX agent

arXiv:2606.00783v1 Announce Type: cross Abstract: Reliable quantification of malaria dynamics in sub-Saharan Africa is hindered by short, noisy, and spatially heterogeneous surveillance records. In Gh

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Before and After Temperature: A Distributional View of Creative LLM Generation

DGX agent

arXiv:2606.01451v1 Announce Type: new Abstract: Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

DGX agent

arXiv:2606.01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differen

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Benchmark Dataset for Catalysis on 2D MXenes

DGX agent

arXiv:2606.00794v1 Announce Type: cross Abstract: Merging first-principles calculations with machine learning (ML), we aim to accelerate the exploration of catalytic behaviour in novel materials. We f

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

DGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

DGX agent

arXiv:2606.01629v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. L

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

DGX agent

arXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

DGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

DGX agent

arXiv:2606.00329v1 Announce Type: cross Abstract: Recursive systems can enter collapse-like regimes -- self-reinforcing amplification, persistent recursion, and narrowing diversity that mask accelerat

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems

DGX agent

arXiv:2606.00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime. This extensibility also creates a supp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset

DGX agent

arXiv:2507.07339v2 Announce Type: replace-cross Abstract: Decisions about managing patients on the heart transplant waitlist are currently made by committees of doctors who consider multiple factors,

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

DGX agent

arXiv:2606.00871v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

DGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA

DGX agent

arXiv:2602.09492v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical ga

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Beyond ell_2-norm and ell_infty-norm: A Curvature-Inspired ell_p-Norm Scheme for Deep Neural Networks

DGX agent

arXiv:2606.02078v1 Announce Type: new Abstract: The existing optimizers for deep neural networks (DNNs) typically rely on either the ell_2 norm or the ell_infty norm, resulting in optimizers that do n

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Beyond Isolated Behaviors: Hierarchical User Modeling for LLM Personalization

DGX agent

arXiv:2606.02300v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet personalizing their outputs to individual users remai

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Beyond Rigid: Benchmarking Non-Rigid Video Editing

DGX agent

arXiv:2601.18340v2 Announce Type: replace Abstract: As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance f

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas

DGX agent

arXiv:2603.19453v2 Announce Type: replace Abstract: We study LLM policy synthesis: using a language model to iteratively generate programmatic agent policies for multi-agent environments. Rather than

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation

DGX agent

arXiv:2508.10312v2 Announce Type: replace Abstract: Recommender systems in concert with Large Language Models (LLMs) present promising avenues for generating semantically-informed recommendations. How

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Beyond Static Gaussians: An Empirical Investigation of Architectural Paradigms for Dynamic 3D Scene Reconstruction

DGX agent

arXiv:2606.00452v1 Announce Type: new Abstract: Dynamic scene reconstruction via 3D Gaussian Splatting (3DGS) has emerged as a compelling approach for representing evolving environments, yet understan

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

DGX agent

arXiv:2606.00065v1 Announce Type: cross Abstract: Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large languag

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Beyond the Simplex: Balanced Prototype Geometry for Scorer-Agnostic Open-Set Recognition

DGX agent

arXiv:2606.01883v1 Announce Type: cross Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical i

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led to 2.2x more code, local agents like original …

DGX agent

Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led to 2.2x more code, local agents like original Claude Code led to 7.4x, & current remote coding agents 17.3

model-releasesethan-mollick--x
2 Jun 2026
Model Releases

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining

DGX agent

arXiv:2510.06048v4 Announce Type: replace Abstract: Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention

DGX agent

arXiv:2512.10414v2 Announce Type: replace Abstract: Recently, reinforcement learning (RL) has become a common choice in enhancing the reasoning capabilities of vision-language models (VLMs). Consideri

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers

DGX agent

arXiv:2606.00957v1 Announce Type: new Abstract: We present a post-training quantization (PTQ) approach for Wan2.1-T2V-14B, a 14-billion-parameter text-to-video diffusion transformer, targeting the W8A

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

BraveGuard: From Open-World Threats to Safer Computer-Use Agents

DGX agent

arXiv:2606.01166v1 Announce Type: cross Abstract: Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shi

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory

DGX agent

arXiv:2606.01385v1 Announce Type: cross Abstract: Software architecture design is a critical yet inherently complex and knowledge-intensive phase that requires balancing competing quality attributes a

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Bridging the Sim-to-Real Gap in Semiconductor Visual Program Synthesis via Input Binarization

DGX agent

arXiv:2606.02434v1 Announce Type: new Abstract: Precise parametric control over circuit geometry is essential for semiconductor inspection, yet obtaining sufficient real training data remains costly.

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Business Utility of Large Language Models as Exploratory Data Analysis Agents

DGX agent

arXiv:2606.00051v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in busines

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

CAFOSat: A Strongly Annotated Dataset for Infrastructure-Aware CAFO Mapping Using High-Resolution Imagery

DGX agent

arXiv:2606.00548v1 Announce Type: cross Abstract: Concentrated Animal Feeding Operations (CAFOs) play an important role in agricultural production but are also associated with environmental, public he

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

DGX agent

arXiv:2505.24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithm

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criter…

DGX agent

Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criteria: they're used both in agent benchmarking and as reward si

model-releasesharrison-chase--x
2 Jun 2026
Model Releases

CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability

DGX agent

arXiv:2606.01495v1 Announce Type: cross Abstract: We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across d

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

DGX agent

arXiv:2606.01995v1 Announce Type: new Abstract: We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language mode

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

CASTLE2026 Team WDL Technical Report

DGX agent

arXiv:2606.00712v1 Announce Type: new Abstract: The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-ch

model-releasesarxiv-cs-cv
2 Jun 2026
← Previous
1…227228229230231…476
Next →