AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
3 Jun 2026

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

Model ReleasesDGX agent

arXiv:2602.05302v3 Announce Type: replace Abstract: We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic

PINNfluence: Interpreting PINNs through Influence Functions

Model ReleasesDGX agent

arXiv:2409.08958v3 Announce Type: replace-cross Abstract: Physics-informed neural networks (PINNs) have emerged as a powerful deep learning approach for solving partial differential equations (PDEs) i

Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.12247v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) present a promising non-sequential paradigm for text generation, distinct from standard autoregressive (AR) a

Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records

Model ReleasesDGX agent

arXiv:2606.02747v1 Announce Type: cross Abstract: Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machin

Pretraining Language Models on Historical Text

Model ReleasesDGX agent

arXiv:2606.02991v1 Announce Type: cross Abstract: We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires add

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

Model ReleasesDGX agent

arXiv:2606.03743v1 Announce Type: new Abstract: While Large Language Models (LLMs) have shown strong performance in generating formal proofs, their outputs often remain less readable, modular, maintai

ProtocolBench: Which LLM MultiAgent Protocol to Choose?

Model ReleasesDGX agent

arXiv:2510.17149v3 Announce Type: replace Abstract: As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and

Psi-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues

Model ReleasesDGX agent

arXiv:2606.02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responde

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

Model ReleasesDGX agent

arXiv:2512.10888v3 Announce Type: replace Abstract: Table extraction (TE) is a key challenge in document understanding. Traditional approaches detect tables first, then recognize their structure. Rece

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03858v1 Announce Type: new Abstract: Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few

q0: Primitives for Hyper-Epoch Pretraining

Model ReleasesDGX agent

arXiv:2606.03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text. But pretraining a single model

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference

Model ReleasesDGX agent

arXiv:2606.02823v1 Announce Type: new Abstract: Two-bit weight quantization is attractive for memory-efficient LLM inference, but the standard W2 level set {-2,-1,0,+1} often collapses under aggressiv

Quadratic integrate-and-fire neurons exhibit less fragmented loss landscapes and outperform leaky integrate-and-fire neurons in spike-based gradient descent

Model ReleasesDGX agent

arXiv:2606.03935v1 Announce Type: cross Abstract: The ability to train spiking neural networks is essential for modeling biological neural networks as well as for neuromorphic computing. However, for

QUIVER: Quantum-Informed Views for Enhanced Representations in Large ML Models

Model ReleasesDGX agent

arXiv:2606.02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example. We introduce QUIVER (Q

Qwen-Image-Flash: Beyond Objective Design

Model ReleasesDGX agent

arXiv:2606.03746v1 Announce Type: cross Abstract: Few-step distillation has become an effective strategy for accelerating advanced visual generative models, yet prior work has largely focused on disti

RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds

Model ReleasesDGX agent

arXiv:2509.18068v2 Announce Type: replace Abstract: Millimeter-wave radar provides robust perception in fog, smoke, dust, and low light, making it attractive for size-, weight-, and power-constrained

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

Model ReleasesDGX agent

arXiv:2606.03728v1 Announce Type: new Abstract: Retrieval-augmented generation systems for legal question answering typically retrieve passages based on semantic similarity and provide them to a langu

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

Model ReleasesDGX agent

arXiv:2606.03889v1 Announce Type: new Abstract: Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real develop

ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models

Model ReleasesDGX agent

arXiv:2510.09711v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and

Reasoning Structure of Large Language Models

Model ReleasesDGX agent

arXiv:2606.03883v1 Announce Type: new Abstract: Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics

Reconstructing Objects along Hand Interaction Timelines in Egocentric Video

Model ReleasesDGX agent

arXiv:2512.07394v2 Announce Type: replace Abstract: We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a

Relational Linearity is a Predictor of Hallucinations

Model ReleasesDGX agent

arXiv:2601.11429v2 Announce Type: replace-cross Abstract: Hallucination is a central failure mode of language models (LMs). We focus on hallucinations in response to questions like: 'Which instrument

Reliability-Guided Depth Fusion for Glare-Resilient Navigation Costmaps

Model ReleasesDGX agent

arXiv:2606.03421v1 Announce Type: new Abstract: Specular glare on reflective floors, glass boundaries, and glossy indoor surfaces frequently corrupts active-stereo RGB-D depth measurements, producing

RESCAST-100K: A Comprehensive Dataset for Cross-Domain Residential Load and Indoor Temperature Forecasting

Model ReleasesDGX agent

arXiv:2606.02852v1 Announce Type: new Abstract: Accurate short-term forecasting of residential energy load and indoor temperature is essential for home energy management systems, grid-level demand res

Rethinking Molecular Text Representations for LLMs: An Empirical Study

Model ReleasesDGX agent

arXiv:2606.03057v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use. We present a sys

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.03784v1 Announce Type: new Abstract: Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning and robotic control, but its effective form and integration strategy remain underexp

RobotValues: Evaluating Household Robots When Human Values Conflict

Model ReleasesDGX agent

arXiv:2606.03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robo

ROBUST-WT: Robust Uncertainty-aware Segmentation Transform via Whitening and Training Enhancements

Model ReleasesDGX agent

arXiv:2606.03069v1 Announce Type: cross Abstract: Generalized segmentation of medical images prevents performance degradation when different imaging devices and clinical protocols are used across mult

RogueMerge: Robust and Unified Attacks against LLM Model Merging

Model ReleasesDGX agent

arXiv:2606.03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a cri

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Model ReleasesDGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series

Model ReleasesDGX agent

arXiv:2606.03301v1 Announce Type: new Abstract: We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize

Sample-Size Scaling of the African Languages NLI Evaluation

Model ReleasesDGX agent

arXiv:2606.03219v1 Announce Type: new Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performan

Samudra 2: Scaling Ocean Emulators across Resolutions

Model ReleasesDGX agent

arXiv:2606.02610v1 Announce Type: cross Abstract: Ocean general circulation models (OGCMs) are essential to climate science but computationally expensive, limiting ensemble size and forcing scenarios.

Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation

Model ReleasesDGX agent

arXiv:2606.03517v1 Announce Type: cross Abstract: Training quantum neural networks (QNNs) on quantum hardware is currently bottlenecked by the cost of gradient estimation: standard parameter-shift met

SCOPE: Real-Time Natural Language Camera Agent at the Edge

Model ReleasesDGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

ScoreStop: Gradient-based early stopping using functional score tests

Model ReleasesDGX agent

arXiv:2606.02740v1 Announce Type: cross Abstract: Gradient boosted decision trees require a stopping rule to avoid overfitting. The standard rule monitors a validation loss and stops if the loss fails

scTranslation: A Comprehensive Benchmark for Single-Cell Multi-Omics Modality Translation

Model ReleasesDGX agent

arXiv:2606.03906v1 Announce Type: new Abstract: Simultaneous measurement of multiple omics modalities in single cells enables researchers to gain a more comprehensive understanding of cellular states

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

Model ReleasesDGX agent

arXiv:2606.03284v1 Announce Type: new Abstract: Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NL

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

Model ReleasesDGX agent

arXiv:2606.03371v1 Announce Type: new Abstract: Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is mad

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

Model ReleasesDGX agent

arXiv:2606.02745v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-speci

SenseJudge: Human-Centric Preference-Driven Judgment Framework

Model ReleasesDGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

shower thought If: 1. AI is smarter than humans at law, therapy, etc. 2. Humans still like talking to other humans. Then: Humans are just an…

Model ReleasesDGX agent

shower thought If: 1. AI is smarter than humans at law, therapy, etc. 2. Humans still like talking to other humans. Then: Humans are just an AI wrapper. Everyone should just regurgitate what Claude te

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Model ReleasesDGX agent

arXiv:2606.03980v1 Announce Type: cross Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) p

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

Model ReleasesDGX agent

arXiv:2606.03056v1 Announce Type: new Abstract: As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend

SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation

Model ReleasesDGX agent

arXiv:2606.03788v1 Announce Type: new Abstract: Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directl

Sources: Benchmark raised 2B across two new funds, including a 1.25B fund for late-stage bets, its first growth fund after decades focusing on new startups (Kate Clark/Wall Street Journal)

Model ReleasesDGX agent

Kate Clark / Wall Street Journal: Sources: Benchmark raised 2B across two new funds, including a 1.25B fund for late-stage bets, its first growth fund after decades focusing on new startups — After a

Sources: DeepSeek is set to raise ~7.4B in its first funding round from investors including Tencent and CATL at a valuation of between ~52B and ~$59B (Reuters)

Model ReleasesDGX agent

Reuters: Sources: DeepSeek is set to raise ~7.4B in its first funding round from investors including Tencent and CATL at a valuation of between ~52B and ~59B — Chinese AI startup DeepSeek is set to ra

Spatial Transcriptomics-Guided Alignment Enhances Molecular Profiling in Pathology Foundation Model

Model ReleasesDGX agent

arXiv:2606.03644v1 Announce Type: new Abstract: Comprehensive molecular profiling is essential for modern precision oncology but remains hindered by prohibitive costs, specimen exhaustion, and protrac

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

Model ReleasesDGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

Startup discovery platform @harmonic_ai rebuilt Scout, their AI platform using Deep Agents and LangSmith. Deep Agents: One frontier model + …

Model ReleasesDGX agent

Startup discovery platform @harmonic_ai rebuilt Scout, their AI platform using Deep Agents and LangSmith. Deep Agents: One frontier model + two tool sets (global company data and firm-specific context

State-Coupled Volatility in Latent Dynamical Systems: Recovery Under Partial Observation

Model ReleasesDGX agent

arXiv:2606.02664v1 Announce Type: cross Abstract: Latent state-space models are widely used to study partially observed dynamical systems, yet most formulations assume that process variability is inde

Staying Alive: Uncensored Survival Analysis with Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2606.03689v1 Announce Type: cross Abstract: Survival Analysis (SA) is a statistical framework that models the time span until some event of interest occurs. Widely used in several domains, inclu

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

Model ReleasesDGX agent

arXiv:2606.02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing be

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

Model ReleasesDGX agent

arXiv:2606.03348v1 Announce Type: cross Abstract: Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic cr

Synthetic Hallucinations, Real Gains: Hard Negatives from Frontier Models for FIM Hallucination Mitigation

Model ReleasesDGX agent

arXiv:2606.03130v1 Announce Type: new Abstract: Small open-source code models that power IDE autocomplete still emit hallucinated Fill-in-the-Middle (FIM) completions: syntactically natural calls to m

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Model ReleasesDGX agent

arXiv:2512.21094v2 Announce Type: replace Abstract: Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet it

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

Model ReleasesDGX agent

arXiv:2606.02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

Model ReleasesDGX agent

arXiv:2509.09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In th

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Model ReleasesDGX agent

arXiv:2606.03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate

← Previous
1…176177178179180…377
Next →