AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
10 Jun 2026

LLM-Based Code Documentation Generation and Multi-Judge Evaluation

Model ReleasesDGX agent

arXiv:2606.09852v1 Announce Type: cross Abstract: High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainabil

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

Model ReleasesDGX agent

arXiv:2606.10956v1 Announce Type: new Abstract: The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade p

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

Model ReleasesDGX agent

arXiv:2606.10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

Model ReleasesDGX agent

arXiv:2606.10768v1 Announce Type: cross Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

Model ReleasesDGX agent

arXiv:2606.10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info ne…

Model ReleasesDGX agent

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info needed to train a frontier model, down to hparams. I also thou

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

Model ReleasesDGX agent

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high-school mathematics, their ability to evaluate the diverse reas

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

Model ReleasesDGX agent

arXiv:2606.09886v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

Model ReleasesDGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Model ReleasesDGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

Model ReleasesDGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

ResearchDGX agent

arXiv:2606.10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals. While recent advances in deep learning have signific

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

Model ReleasesDGX agent

arXiv:2606.09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once. Clinicians must determine whether a patient should receive tar

wooh https://x.com/shadcn/status/2064671802509410806?s=46

Model ReleasesDGX agent

wooh https://x.com/shadcn/status/2064671802509410806?s=46 You have Claude Fable for only a few days. Here's how to make the most of it. Introducing /improve: use your most capable model to audit your

9 Jun 2026

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

Model ReleasesDGX agent

arXiv:2606.08669v1 Announce Type: cross Abstract: Voice biometric systems face growing threats from spoofing attacks, yet the evaluation of detection models remains inconsistent across datasets. To in

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks

SafetyDGX agent

arXiv:2604.01039v2 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP

Model ReleasesDGX agent

arXiv:2505.11189v3 Announce Type: replace Abstract: Large language models (LLMs) can amplify misinformation, undermining societal goals such as the UN SDGs. We study three documented drivers of misinf

Capacity, Not Format: Rethinking Structured Reasoning Failures

ResearchDGX agent

arXiv:2606.09410v1 Announce Type: new Abstract: Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capac

CATPO: Critique-Augmented Tree Policy Optimization

Model ReleasesDGX agent

arXiv:2606.08346v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

Model ReleasesDGX agent

arXiv:2606.08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observabili

CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

Model ReleasesDGX agent

arXiv:2606.08169v1 Announce Type: cross Abstract: Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation model

Claude Fable 5 now available on AI Gateway

Model ReleasesDGX agent

Claude Fable 5 is now available through Vercel's AI Gateway, expanding the model options developers can access through the platform. This announcement likely details how developers can integrate and u

CTS-Bench: Benchmarking Graph Coarsening Trade-offs for GNNs in Clock Tree Synthesis

Model ReleasesDGX agent

arXiv:2602.19330v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) are increasingly explored for physical design analysis in Electronic Design Automation, particularly for modeling Clock

Curvature-Guided LoRA: Matching Full Fine-Tuning in Function Space

Model ReleasesDGX agent

arXiv:2603.29824v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning i

Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan

Model ReleasesDGX agent

arXiv:2606.09767v1 Announce Type: cross Abstract: Neural machine translation for digitally low-resource Indigenous languages is often hindered by extreme data scarcity, prompting reliance on extractiv

De novo molecular generation with optical property preconditioning at the token level

Model ReleasesDGX agent

arXiv:2606.08221v1 Announce Type: new Abstract: Designing OLED molecules with targeted optical properties remains challenging due to the scarcity of high-quality data and the limited reliability of co

Diffuse AI Control on Fuzzy Tasks

SafetyDGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

Enhancing AI Interpretability and Safety through Localised Architectures

SafetyDGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA

Model ReleasesDGX agent

arXiv:2606.07548v1 Announce Type: cross Abstract: The MedHopQA challenge presents a critical test for Large Language Models (LLMs): complex, multi-hop reasoning in the high-stakes biomedical domain. T

Explaining Data Mixing Scaling Laws

TutorialsDGX agent

arXiv:2606.08167v1 Announce Type: cross Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understandin

Few-step Cofolding with All-Atom Flow Maps

SafetyDGX agent

arXiv:2606.08375v1 Announce Type: new Abstract: All-atom generative modeling of 3D biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-liga

Fluid, natural voice translation with Gemini 3.5 Live Translate

Model ReleasesDGX agent

Gemini 3.5 Live Translate is an audio model delivering near real-time speech-to-speech translation in over 70 languages , with automatic language detection and natural-sounding translated speech that

Forecasting Japanese elections: A nonlinear machine-learning approach

ResearchDGX agent

arXiv:2606.07572v1 Announce Type: cross Abstract: Despite Japan being one of the world's largest advanced democracies, the development of election forecasting models for its national elections remains

From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing

Model ReleasesDGX agent

arXiv:2606.08932v1 Announce Type: cross Abstract: Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but si

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background s

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

SafetyDGX agent

arXiv:2606.08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-wor

Harnessing Streaming Video in the Wild

Model ReleasesDGX agent

arXiv:2606.08615v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentar

Integrating Deep Learning Demand Forecasting with Multi-Objective Optimization for Circular Coffee Supply Chains: A Data-Driven Framework for Cost, Emissions, and Freshness Management

Model ReleasesDGX agent

arXiv:2606.08314v1 Announce Type: new Abstract: The coffee supply chain is one of the most complex agri-food networks, marked by geographically dispersed production, multi-tier coordination, and high

Integrating gene regulatory priors into Transformer attention with scTransformer for interpretable scRNA-seq analysis

ResearchDGX agent

arXiv:2606.09558v1 Announce Type: cross Abstract: Motivation: Transformer-based models are increasingly applied to large-scale single-cell transcriptomics, showing strong performance through self-supe

Internalizing Geometric Law: Learning from Solver Residuals for Precision-Critical Generation

Model ReleasesDGX agent

arXiv:2606.09278v1 Announce Type: cross Abstract: Large Language Models frequently hallucinate in precision-critical domains such as technical diagramming and mechanical design, where outputs must sat

Introducing the Fast Gemma Challenge with Hugging Face Over the next few days, dozens of agents will collaborate to make Gemma 4 E4B even fa…

Model ReleasesDGX agent

Google and Hugging Face are launching the Fast Gemma Challenge, where multiple agents will collaborate to optimize the performance and speed of Gemma 4 E4B model. The initiative aims to improve the ef

KPGrasp: Scalable Keypoint Flow Matching for Dexterous Grasp Generation

Model ReleasesDGX agent

arXiv:2606.09314v1 Announce Type: new Abstract: Generating high-quality dexterous grasps remains challenging for learning-based methods, which often depend on carefully tuned contact losses or costly

Learning Predictive Control with Deep Koopman Operators for Autonomous Vehicle Motion Planning

Model ReleasesDGX agent

arXiv:2606.08136v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used for autonomous-vehicle (AV) motion planning, but its real-time applicability is often limited by the need

Learning to Solve Generative ODEs Beyond the Linear Span

ResearchDGX agent

arXiv:2606.08672v1 Announce Type: new Abstract: Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver lear

LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load

Model ReleasesDGX agent

arXiv:2603.23640v2 Announce Type: replace-cross Abstract: Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

Model ReleasesDGX agent

arXiv:2601.22859v3 Announce Type: replace-cross Abstract: The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bot

Minibatch Selection via Partition Matroid Constrained Gradient Matching

Model ReleasesDGX agent

arXiv:2606.07954v1 Announce Type: cross Abstract: Training large language models (LLMs) on heterogeneous data requires selecting minibatches that balance convergence speed with coverage across domains

OmniGen-AR: AutoRegressive Any-to-Image Generation

Model ReleasesDGX agent

arXiv:2606.09156v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated strong potential in visual generation, offering superior performance with simple architectures and optimiza

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

Model ReleasesDGX agent

arXiv:2606.07577v1 Announce Type: new Abstract: Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited

Partially Performative Prediction

ResearchDGX agent

arXiv:2606.07890v1 Announce Type: new Abstract: Performative prediction studies feedback loops that arise when predictive models are deployed in consequential domains. In these settings, deploying a m

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

Model ReleasesDGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

phepy: Visual benchmarks and improvements for out-of-distribution detectors

Model ReleasesDGX agent

arXiv:2503.05169v2 Announce Type: replace Abstract: Applying machine learning to increasingly high-dimensional problems with sparse or biased training data increases the risk that a model is used on i

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects

Model ReleasesDGX agent

arXiv:2606.08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can beha

Read the blog to learn more: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/

Model ReleasesDGX agent

This blog post from Google AI announces features and updates related to Gemini models, likely covering new capabilities for Gemini Live, version 3.5, and translation functionality. The announcement de

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

Model ReleasesDGX agent

arXiv:2601.06649v2 Announce Type: replace-cross Abstract: Research in machine learning has questioned whether increases in training token counts reliably produce proportional performance gains in larg

SC3: The Multi-Solvent Solubility Challenge and Benchmark

Model ReleasesDGX agent

arXiv:2606.07656v1 Announce Type: cross Abstract: Solubility prediction is a standard benchmark in computational chemistry, yet multi-solvent models which reportedly approach the experimental-noise ce

SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

Model ReleasesDGX agent

arXiv:2602.09809v2 Announce Type: replace Abstract: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorr

Self-Consistent Generative Paths via Admissible Random Variational Transport

Local AiDGX agent

arXiv:2606.08953v1 Announce Type: new Abstract: Modern generative models often define an entire probability path from a simple prior to the data law, rather than only an endpoint map. Diffusion models

← Previous
1…342343344345346…1042
Next →