AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlog
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,794 results
Model Releases

HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

DGX agent

arXiv:2605.28315v1 Announce Type: new Abstract: General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language

model-releasesarxiv-cs-cl
28 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Integrated and Cross-Architecture Interpretation of LLM Reasoning

DGX agent

arXiv:2605.28006v1 Announce Type: cross Abstract: Understanding how LLMs reason is hindered by a practical asymmetry: while their generated outputs are observable, the underlying reasoning patterns re

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

DGX agent

arXiv:2605.28714v1 Announce Type: cross Abstract: An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking

DGX agent

arXiv:2605.27914v1 Announce Type: cross Abstract: Subjective evaluation of LLM behavior -- empathy, restraint, calibrated emotional tone -- is hard. Human inter-rater agreement on such qualities satur

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

DGX agent

arXiv:2605.28721v1 Announce Type: new Abstract: Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diag

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

DGX agent

arXiv:2603.21165v2 Announce Type: replace Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresente

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs

DGX agent

arXiv:2510.01724v2 Announce Type: replace Abstract: Mass spectrometry-based metabolomics generates complex, high-dimensional data that holds vast potential for biological discovery but remains difficu

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding

DGX agent

arXiv:2505.21771v2 Announce Type: replace-cross Abstract: Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remai

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

DGX agent

arXiv:2605.27853v1 Announce Type: new Abstract: We present MolLingo, a multi-agent system that emulates the reasoning process of a chemist to automate molecular design. Existing LLM-based approaches e

model-releasesarxiv-cs-ai
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Model Releases

Not All Pixels Are Equal: Pixel-wise Meta-Learning for Medical Segmentation with Noisy Labels

DGX agent

arXiv:2511.18894v5 Announce Type: replace-cross Abstract: Medical image segmentation is crucial for clinical applications, but it is frequently disrupted by noisy annotations and ambiguous anatomical

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

On the Equivariant Learning of the Q-tensor Order Parameter

DGX agent

arXiv:2605.27679v1 Announce Type: cross Abstract: We construct and evaluate group-equivariant neural networks for the prediction of the two-dimensional Q-tensor order parameter of nematic liquid cryst

model-releasesarxiv-cs-cv
28 May 2026
Agents

Opus 4.8 is now supported in Hermes Agent ^_^

DGX agent

Nous Research has announced support for Opus 4.8 in the Hermes Agent framework. This update enables the Hermes Agent to utilize Anthropic's Opus 4.8 model, expanding the available model options for us

agentsnous-research--x
28 May 2026
Model Releases

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

DGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

DGX agent

arXiv:2605.27887v1 Announce Type: new Abstract: LLMs have shown strong performance across diverse financial tasks, yet portfolio management (PM), a critical financial decision-making task, remains poo

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

DGX agent

arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage

model-releasesarxiv-cs-cl
28 May 2026
Safety

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

DGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

safetyarxiv-cs-cl
28 May 2026
Model Releases

Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit

DGX agent

arXiv:2605.27439v1 Announce Type: cross Abstract: AI assistants like ChatGPT and Claude are recommendation engines, not search engines: they answer commercial queries by directly nominating brands rat

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text

DGX agent

arXiv:2605.28363v1 Announce Type: new Abstract: Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations,

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

DGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

DGX agent

arXiv:2603.05642v2 Announce Type: replace-cross Abstract: Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surroun

model-releasesarxiv-cs-ai
28 May 2026
Research

Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation

DGX agent

arXiv:2605.27993v1 Announce Type: new Abstract: Object hallucination remains a primary obstacle to the reliable deployment of Multimodal Large Language Models (MLLMs). Current inference-time mitigatio

researcharxiv-cs-cl
28 May 2026
Model Releases

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

DGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving

DGX agent

arXiv:2605.28136v1 Announce Type: new Abstract: Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (

model-releasesarxiv-cs-cv
28 May 2026
Research

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

DGX agent

arXiv:2605.28741v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of

researcharxiv-cs-cv
28 May 2026
Model Releases

StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment

DGX agent

arXiv:2605.28073v1 Announce Type: cross Abstract: Story rewriting aims to adapt existing narratives to diverse reader preferences while preserving plot consistency and narrative coherence. Unlike conv

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

DGX agent

arXiv:2605.27431v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across dive

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

DGX agent

arXiv:2605.28063v1 Announce Type: cross Abstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. C

model-releasesarxiv-cs-ai
28 May 2026
Safety

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

DGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

safetyarxiv-cs-lg
28 May 2026
Model Releases

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

DGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

DGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Weak Convergence Analysis of Online Neural Actor-Critic Algorithms

DGX agent

arXiv:2403.16825v2 Announce Type: replace Abstract: We prove that a single-layer neural network trained with the online actor critic algorithm converges in distribution to a random ordinary differenti

model-releasesarxiv-cs-lg
28 May 2026
Safety

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

DGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

safetyarxiv-cs-ai
28 May 2026
Model Releases

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

DGX agent

arXiv:2605.27586v1 Announce Type: cross Abstract: Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. W

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

DGX agent

arXiv:2605.27116v1 Announce Type: new Abstract: Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-wor

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Curation and Extraction of Drug-Related Entities from Reddit Platform

DGX agent

arXiv:2605.26445v1 Announce Type: new Abstract: Physicians learn primarily about illicit drugs from clinical overdose cases, limiting their understanding of real-world usage. Meanwhile, drug users sha

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

DGLD: Domain-Gated Latent Diffusion for the Discovery of Novel Energetic Materials

DGX agent

arXiv:2605.26540v1 Announce Type: cross Abstract: Energetic-materials performance gains translate directly into reduced propellant mass, smaller warheads, and more efficient civilian gas-generators, y

model-releasesarxiv-cs-ai
27 May 2026
Applications

Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?

DGX agent

arXiv:2605.27135v1 Announce Type: cross Abstract: With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI

applicationsarxiv-cs-cv
27 May 2026
Model Releases

FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning

DGX agent

arXiv:2603.13282v2 Announce Type: replace-cross Abstract: Federated Learning (FL) with Low-Rank Adaptation (LoRA) has become a standard for privacy-preserving LLM fine-tuning. However, existing person

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

DGX agent

arXiv:2605.26579v1 Announce Type: new Abstract: The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement lea

model-releasesarxiv-cs-lg
27 May 2026
Tutorials

GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision

DGX agent

arXiv:2603.09551v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reason

tutorialsarxiv-cs-cv
27 May 2026
Agents

Hermes Agent now has a built-in MCP Catalog

DGX agent

Nous Research announced that their Hermes Agent now includes an integrated MCP (Model Context Protocol) Catalog, enabling users to discover and access available Model Context Protocol tools and integr

agentsnous-research--x
27 May 2026
Model Releases

If we had done everything I suggested in my 2020 arXiv article “The Next Decade in AI”, we might actually have reached AGI by now. In the la…

DGX agent

If we had done everything I suggested in my 2020 arXiv article “The Next Decade in AI”, we might actually have reached AGI by now. In the last three years, after a detour driven by the false promise o

model-releasesgary-marcus--x
27 May 2026
Research

Inferring Group Intent as a Cooperative Game. An NLP-based Framework for Trajectory Analysis

DGX agent

arXiv:2510.23905v2 Announce Type: replace-cross Abstract: This paper studies group target trajectory intent as the outcome of a cooperative game where the complex-spatio trajectories are modeled using

researcharxiv-cs-lg
27 May 2026
Model Releases

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

DGX agent

arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes

model-releasesarxiv-cs-ai
27 May 2026
Research

LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems

DGX agent

arXiv:2512.01556v3 Announce Type: replace Abstract: Foundation models often generate unreliable answers, while heuristic uncertainty estimators fail to fully distinguish correct from incorrect outputs

researcharxiv-cs-ai
27 May 2026
Research

Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

DGX agent

arXiv:2605.27268v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. W

researcharxiv-cs-ai
27 May 2026
Model Releases

MerLean-Prover: A Recursive Looping Harness for End-to-End Lean 4 Theorem Proving

DGX agent

arXiv:2605.26959v1 Announce Type: cross Abstract: MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types

model-releasesarxiv-cs-cl
27 May 2026
← Previous
1…607608609610611…1371
Next →