AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,593 results
Model Releases

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests

DGX agent

arXiv:2607.00990v1 Announce Type: cross Abstract: Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue re

model-releasesarxiv-cs-ai
2 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives

DGX agent

arXiv:2606.16511v2 Announce Type: replace Abstract: Recent work motivates moving large language model (LLM) evaluation from mean-based to tail-aware metrics, including conditional value-at-risk and ta

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

TallyTrain: Communication-Efficient Federated Distillation

DGX agent

arXiv:2607.00173v1 Announce Type: new Abstract: Federated learning is bandwidth-bound on two orthogonal axes: model size, which limits how often parameter-averaging methods can afford to merge, and cl

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification

DGX agent

arXiv:2508.17519v3 Announce Type: replace-cross Abstract: Handling missing data in time series classification remains a significant challenge in various domains. Traditional methods often rely on impu

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

DGX agent

arXiv:2510.10180v2 Announce Type: replace Abstract: Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial v

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

DGX agent

arXiv:2606.13148v2 Announce Type: replace Abstract: Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite im

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds

DGX agent

arXiv:2607.00276v1 Announce Type: cross Abstract: Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

The Download: a startup has a solution for AI’s groupthink problem

DGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. LLMs are stuck in a groupthink groove. This startup is trying

model-releasesmit-tech-review
2 Jul 2026
Model Releases

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

DGX agent

arXiv:2607.01033v1 Announce Type: new Abstract: Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

The talk about Mythos and cybersecurity was not, in fact, hype. (As anyone using Fable to do autonomous work has probably recognized)

DGX agent

The talk about Mythos and cybersecurity was not, in fact, hype. (As anyone using Fable to do autonomous work has probably recognized) AI appears to be finding software vulnerabilities at scale. In Jun

model-releasesethan-mollick--x
2 Jul 2026
Model Releases

Timesynth: A Temporal Fidelity Framework for Health Signal Digital Twins

DGX agent

arXiv:2607.00431v1 Announce Type: new Abstract: Forecasting models for health-signal digital twins must preserve the oscillatory, frequency, phase, and state-transition dynamics of physiological signa

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

Toward Cybersecurity-Expert Small Language Models

DGX agent

arXiv:2510.14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domai

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

DGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Towards Metric-Agnostic Trajectory Forecasting

DGX agent

arXiv:2607.01133v1 Announce Type: new Abstract: Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavio

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

TRIE: An Evaluation Framework for Stochastic PDE Surrogates

DGX agent

arXiv:2607.00196v1 Announce Type: new Abstract: Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

DGX agent

arXiv:2511.18050v1 Announce Type: cross Abstract: Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K acro

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics

DGX agent

arXiv:2607.00969v1 Announce Type: cross Abstract: Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approach

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

DGX agent

arXiv:2601.04453v4 Announce Type: replace Abstract: World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recen

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Using DSPy to evaluate and improve Datasette Agent's SQL system prompts

DGX agent

Research: Using DSPy to evaluate and improve Datasette Agent's SQL system prompts One of this morning's AIE keynotes covered dspy, which reminded me I've been meaning to see if it could help me improv

model-releasessimon-willison
2 Jul 2026
Model Releases

Validating Causal Abstraction Metrics on Simulated Complex Systems

DGX agent

arXiv:2607.00267v1 Announce Type: cross Abstract: A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lowe

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

DGX agent

arXiv:2503.13445v3 Announce Type: replace-cross Abstract: When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithful,

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

DGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning

DGX agent

arXiv:2603.17720v2 Announce Type: replace Abstract: Imitation learning is a prominent paradigm for robotic manipulation. However, existing visual imitation methods map 2D image observations directly t

model-releasesarxiv-cs-ro
2 Jul 2026
Model Releases

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

DGX agent

arXiv:2607.00302v1 Announce Type: new Abstract: Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

We are hiring our founding team in Korea 🇰🇷 Join us! P.S. Mistral will be at @icmlconf (July 6–11). Come meet the team!

DGX agent

Mistral AI is recruiting for its founding team in Korea and will have representatives attending ICML conference from July 6-11, 2024, where interested candidates can meet the team in person.

model-releasesarthur-mensch--x
2 Jul 2026
Model Releases

What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models

DGX agent

arXiv:2607.00283v1 Announce Type: cross Abstract: Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat a

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps

DGX agent

arXiv:2607.00004v1 Announce Type: cross Abstract: While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the agi

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Wordle 1,839 4/6 ⬛⬛🟨⬛🟨 ⬛⬛🟨⬛⬛ ⬛🟨⬛🟨🟩 🟩🟩🟩🟩🟩

DGX agent

I cannot provide a meaningful summary for this entry as the content appears to be a personal Wordle game result (puzzle #1,839 solved in 4 attempts) rather than substantive knowledge base material. Th

model-releasesanthropic--x
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

DGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

DGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Op…

DGX agent

You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Opus 4.8. (This is one reason why I am skeptical of just swapp

model-releasesethan-mollick--x
2 Jul 2026
Model Releases

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in d…

DGX agent

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in different formats. The second your team uses more than one (t

model-releasesharrison-chase--x
2 Jul 2026
Model Releases

Z.ai launches ZCode, an 'Agentic Development Environment' optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month (Michael Nuñez/VentureBeat)

DGX agent

Michael Nuñez / VentureBeat: Z.ai launches ZCode, an “Agentic Development Environment” optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month — The move marks th

model-releasestechmeme
2 Jul 2026
Model Releases

ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces

DGX agent

arXiv:2607.01125v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables fine-tuning large language models when backpropagation is unavailable or memory-prohibitive, but existing methods

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

A Large-Language-Model Supported Personalized Driving Framework for Lane Change in Highway Scenarios

DGX agent

arXiv:2606.31483v1 Announce Type: new Abstract: Personalized driving can improve the user acceptance of automated driving systems. However, existing methods still provide limited support for translati

model-releasesarxiv-cs-ro
1 Jul 2026
Model Releases

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

DGX agent

arXiv:2404.10034v3 Announce Type: replace Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

A Reproducible Benchmark of Lightweight CNNs: Accuracy, Efficiency, and the Impact of Pretrained Initialization

DGX agent

arXiv:2505.03303v3 Announce Type: replace-cross Abstract: Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pretr

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

DGX agent

arXiv:2606.31763v1 Announce Type: new Abstract: Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experi

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

DGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection

DGX agent

arXiv:2606.30837v1 Announce Type: cross Abstract: The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases training a

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A swap-adversarial framework for improving domain generalization in electrocorticography-based Parkinson's disease classification

DGX agent

arXiv:2602.10528v2 Announce Type: replace-cross Abstract: We propose a novel swap-adversarial framework that mitigates high inter-subject variability and the high-dimensional low-sample-size problem i

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control

DGX agent

arXiv:2606.30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones. One reason

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management

DGX agent

arXiv:2606.30997v1 Announce Type: new Abstract: We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior f

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A time-series classification framework for individual-level absenteeism prediction under severe class imbalance

DGX agent

arXiv:2606.31532v1 Announce Type: new Abstract: Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat processing, constr

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

A Transferable Learned Temporal Prior for Transmission Reconstruction and Decision-Relevant Uncertainty in Real Outbreak Labels

DGX agent

arXiv:2606.30842v1 Announce Type: new Abstract: Outbreak transmission reconstruction treats epidemiological timing and transmission labels as deterministic ground truth; neither has been systematicall

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Absorption-Feature-Guided Distance-Decoupled Estimation and Band Selection for LWIR Hyperspectral Passive Ranging

DGX agent

arXiv:2606.31824v1 Announce Type: new Abstract: Long-wave infrared (LWIR) hyperspectral observations contain distance-dependent atmospheric absorption signatures, providing a physical basis for long-r

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification

DGX agent

arXiv:2606.30702v1 Announce Type: cross Abstract: Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey sampling, demog

model-releasesarxiv-cs-ai
1 Jul 2026
← Previous
1…140141142143144…471
Next →