AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
2 Jun 2026

Data agents don't fail at writing SQL. They fail at knowing your business. Schemas show you the columns, but they don't tell you which view …

Model ReleasesDGX agent

Data agents don't fail at writing SQL. They fail at knowing your business. Schemas show you the columns, but they don't tell you which view is canonical for ARR, how often each metric updates, or whic

Data Collection for Training Quality-Control AI in Carpet Manufacturing

Model ReleasesDGX agent

arXiv:2606.01023v1 Announce Type: cross Abstract: Visual inspection remains the dominant quality-control practice in woven and tufted carpet production, yet it is slow, subjective, and inconsistent at

datasette-agent-micropython 0.1a0

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Release: datasette-agent-micropython 0.1a0 I want Datasette Agent to be able to generate and execute Python code safely. This alpha is looking promising so far. GPT-5.5 has so far failed to break out

Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging

Model ReleasesDGX agent

arXiv:2606.01717v1 Announce Type: new Abstract: Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback

Model ReleasesDGX agent

arXiv:2606.01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy. For conte

DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations

Model ReleasesDGX agent

arXiv:2606.02289v1 Announce Type: new Abstract: Existing hallucination taxonomies classify LLM errors by what is wrong with the output -- memorised misconceptions, reasoning failures, fluent fabricati

Deep Research as Rubric for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.01091v1 Announce Type: new Abstract: Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a p

Deformable Wiener Filter for Future Video Coding

Model ReleasesDGX agent

arXiv:2606.01576v1 Announce Type: new Abstract: In-loop filters have attracted increasing attention due to the remarkable noise-reduction capability in the hybrid video coding framework. However, the

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

Model ReleasesDGX agent

arXiv:2606.01710v1 Announce Type: new Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlat

Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design

Model ReleasesDGX agent

arXiv:2603.13312v2 Announce Type: replace-cross Abstract: Interior design is a requirements-to-visual-plan generation process that must simultaneously satisfy verifiable spatial feasibility and compar

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

Model ReleasesDGX agent

arXiv:2505.16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the lo

Diagnosing LLM Arbitration Behavior over Pre-evidence Epistemic States in RAG-based Fact-Checking

Model ReleasesDGX agent

arXiv:2606.01120v1 Announce Type: new Abstract: In RAG-based fact-checking, LLMs are increasingly used as verifiers to check given claims against retrieved evidence. Their parametric knowledge can ind

Differentially Private Datastore Generation for Retrieval-Augmented Inference

Model ReleasesDGX agent

arXiv:2606.01413v1 Announce Type: cross Abstract: It is crucial for modern on-device AI systems that rely on retrieval-augmented inference to release and share datastores without compromising individu

DINO-GFSA: Geo-Localization via Semantic Gated Fusion and Mamba-based Sequential Aggregation

Model ReleasesDGX agent

arXiv:2606.00784v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) is critical for Unmanned Aerial Vehicle (UAV) self-positioning and target localization in GNSS-denied environments. H

Disentanglement-Based Equivariant Learning for Compositional VQA

Model ReleasesDGX agent

arXiv:2606.02168v1 Announce Type: new Abstract: Compositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of pre

Disentangling Similarity and Relatedness in Topic Models

Model ReleasesDGX agent

arXiv:2603.10619v2 Announce Type: replace Abstract: The recent success of large pre-trained language models (PLMs) has motivated their integration into topic modeling. However, PLM-augmented topic mod

Distillation of Large Language Models via Concrete Score Matching

Model ReleasesDGX agent

arXiv:2509.25837v3 Announce Type: replace-cross Abstract: Large language models (LLMs) deliver remarkable performance but are costly to deploy, motivating knowledge distillation (KD) for efficient inf

DLLM-JEPA: Joint Embedding Predictive Architectures for Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2606.00091v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) have reshaped self-supervised representation learning in vision. The recent LLM-JEPA ported JEPA to a

Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark

Model ReleasesDGX agent

arXiv:2606.02214v1 Announce Type: new Abstract: Large language models are increasingly used in value-sensitive decision settings, where irrelevant demographic cues should not alter judgments. We const

Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

Model ReleasesDGX agent

arXiv:2606.02357v1 Announce Type: cross Abstract: Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interp

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

Model ReleasesDGX agent

arXiv:2606.00477v1 Announce Type: new Abstract: Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world ap

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

Model ReleasesDGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

Domain-Shift-Aware Conformal Prediction for Large Language Models

Model ReleasesDGX agent

arXiv:2510.05566v2 Announce Type: replace-cross Abstract: Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factuall

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Model ReleasesDGX agent

arXiv:2606.01393v1 Announce Type: cross Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Opt

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

Model ReleasesDGX agent

arXiv:2606.01434v1 Announce Type: new Abstract: Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each ci

Dynamic Proxy-Mixing: Transferring Replay Controllers from Small to Large Models for Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2606.00400v1 Announce Type: new Abstract: Continual instruction tuning updates a language model through a sequence of new domains, yet each update can progressively erode previously learned capa

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

Model ReleasesDGX agent

arXiv:2606.02280v1 Announce Type: new Abstract: Real-world dynamics shifts pose a critical challenge for reinforcement learning in robotics, as policies tightly coupled to nominal environments often f

Early Prediction of Liver Cirrhosis Up to Two Years in Advance: A Machine Learning Study Benchmarking Against the FIB-4 and APRI Scores

Model ReleasesDGX agent

arXiv:2601.00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

Model ReleasesDGX agent

arXiv:2606.01909v1 Announce Type: cross Abstract: We present Echo, a proof-of-concept audio system built around a single 25 M-parameter ViT encoder. The encoder is pretrained with a JEPA objective and

Echo State Networks for Time Series Forecasting: Hyperparameter Sweep and Benchmarking

Model ReleasesDGX agent

arXiv:2602.03912v4 Announce Type: replace Abstract: This paper investigates the performance of Echo State Networks (ESNs) for univariate forecasting of monthly and quarterly time series from the M4 Fo

Efficient Exploration for Iterative Nash Preference Optimization

Model ReleasesDGX agent

arXiv:2606.01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences a

Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking

Model ReleasesDGX agent

arXiv:2606.01240v1 Announce Type: new Abstract: The demand for powerful instruction following and reasoning capability of large language models (LLMs) has promoted rapid development of retrieval-augme

Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark

Model ReleasesDGX agent

arXiv:2606.02246v1 Announce Type: new Abstract: To operate in the physical world, embodied agents must perceive their environment in an 'always-on' fashion, selectively accessing the most informative

Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification

Model ReleasesDGX agent

arXiv:2412.03771v3 Announce Type: replace-cross Abstract: Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and te

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

Model ReleasesDGX agent

arXiv:2606.00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can im

Empathy Applicability Modeling for General Health Queries

Model ReleasesDGX agent

arXiv:2601.09696v2 Announce Type: replace Abstract: LLMs are increasingly being integrated into clinical workflows, yet they often lack clinical empathy, an essential aspect of effective doctor-patien

Enhancing Blind Source Separation with Dissociative Principal Component Analysis

Model ReleasesDGX agent

arXiv:2411.12321v2 Announce Type: replace Abstract: Principal component analysis (PCA) and its sparse variants (sPCA) are widely used as a precursor to independent component analysis (ICA) for blind s

Enhancing LLM Metacognition via Cognitive Pairwise Training

Model ReleasesDGX agent

arXiv:2606.00869v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become central to LLM reasoning, but its outcome-level rewards can make models more willing to

Error Bounds for a Diffusion Model-Based Drift Estimator

Model ReleasesDGX agent

arXiv:2606.02115v1 Announce Type: cross Abstract: Parameter estimation in stochastic differential equations is a classical statistical problem of much importance in many scientific fields. Recent work

ES-Merging: Biological MLLM Merging via Embedding Space Signals

Model ReleasesDGX agent

arXiv:2603.14405v2 Announce Type: replace-cross Abstract: Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing mod

Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding

Model ReleasesDGX agent

arXiv:2603.03312v3 Announce Type: replace-cross Abstract: Decoding natural language from non-invasive EEG signals is a promising yet challenging task. However, current state-of-the-art models remain c

Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization

Model ReleasesDGX agent

arXiv:2606.00544v1 Announce Type: cross Abstract: Modern language-model fine-tuning typically pairs each prompt with a single response, even though many prompts admit multiple valid completions. This

EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams

Model ReleasesDGX agent

arXiv:2603.27223v2 Announce Type: replace-cross Abstract: We present EuraGovExam, a multilingual and multimodal benchmark sourced from real-world civil service examinations across five representative

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games

Model ReleasesDGX agent

arXiv:2606.00103v1 Announce Type: new Abstract: We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Where

Evaluating Real-World Generalizability of Algorithm Selection Models

Model ReleasesDGX agent

arXiv:2606.02016v1 Announce Type: new Abstract: Algorithm Selection (AS) aims to automatically identify the most suitable optimization algorithm for a given problem instance by leveraging measurable p

Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers

Model ReleasesDGX agent

arXiv:2602.22221v2 Announce Type: replace-cross Abstract: Search engines and AI-powered systems increasingly mediate access to factual information, yet their reliability remains difficult to evaluate

Evaluating the Reversal Curse in Model Editing

Model ReleasesDGX agent

arXiv:2310.10322v3 Announce Type: replace Abstract: Large language models (LLMs) are prone to hallucinate unintended text due to false or outdated knowledge. Since retraining LLMs is resource intensiv

Experimenting with TPUs, GKE Managed DRANET, and Multi-cluster Inference Gateway

Model ReleasesDGX agent

What happens when your workload fails in one region but you need access to service? This is a common case for availability and uptime. With recent enhancement to the Kubernetes ecosystem and capabilit

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

Model ReleasesDGX agent

arXiv:2606.02402v1 Announce Type: new Abstract: The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through

Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression

Model ReleasesDGX agent

arXiv:2606.01608v1 Announce Type: new Abstract: Most existing extreme compression methods fail to achieve an optimal rate-distortion-perception trade-off, as they typically prioritize perceptual fidel

Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays

Model ReleasesDGX agent

arXiv:2509.15234v2 Announce Type: replace Abstract: Multimodal learning from paired medical images and clinical text is a central challenge in medical data-driven informatics, where effective cross-mo

ExpWeaver: LLM Agents Learn from Experience via Latent RAG

Model ReleasesDGX agent

arXiv:2606.01041v1 Announce Type: new Abstract: Experience learning has achieved promising results in enhancing LLM agent planning and reasoning by integrating past interactions as reusable knowledge.

FACT: A Simple and Efficient Framework for Active Finetuning

Model ReleasesDGX agent

arXiv:2606.02079v1 Announce Type: new Abstract: The main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected

FALAT: Tracing Failures in LLM Agent Trajectories via Dependency-Guided Search

Model ReleasesDGX agent

arXiv:2606.00765v1 Announce Type: new Abstract: LLM-based agents increasingly solve complex tasks through long trajectories involving reasoning steps, tool calls, and inter-agent communication. Howeve

Fast-SAM3D: 3Dfy Anything in Images but Faster

Model ReleasesDGX agent

arXiv:2602.05293v2 Announce Type: replace Abstract: SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this w

Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing

Model ReleasesDGX agent

arXiv:2606.02218v1 Announce Type: cross Abstract: Synchronous reinforcement learning methods such as Group Relative Policy Optimization (GRPO) provide stable and reproducible on-policy training, but t

Feature to Dynamics: Feature-space to Autoregression strategy for Zero-shot Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.01289v1 Announce Type: new Abstract: Zero-shot time series forecasting aims to predict future values for previously unseen series, requiring models to generalize temporal dynamics beyond th

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

Model ReleasesDGX agent

arXiv:2604.03893v2 Announce Type: replace Abstract: Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and th

FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes

Model ReleasesDGX agent

arXiv:2606.02523v1 Announce Type: new Abstract: Suicide memes are memes used to express suicide-related thoughts or comment on suicide-related issues. Suicide memes are increasingly common on social m

Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models

Model ReleasesDGX agent

arXiv:2504.03635v4 Announce Type: replace Abstract: Reasoning is a core capability of language models (LMs), yet it remains unclear how much model capacity is necessary to support reasoning during pre

← Previous
1…179180181182183…377
Next →