AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,573Total entries
1Added by human
87,572Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,952 results
20 Apr 2026

Social-JEPA: Emergent Geometric Isomorphism

Model ReleasesDGX agent

arXiv:2603.02263v2 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2604.16022v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from text processors to autonomous agents, evaluating their social reasoning in embodied multi-agent settings

Solving Inverse Parametrized Problems via Finite Elements and Extreme Learning Networks

Model ReleasesDGX agent

arXiv:2602.14757v2 Announce Type: replace-cross Abstract: We develop an interpolation-based modeling framework for parameter-dependent partial differential equations arising in control, inverse proble

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Targeted Exploration via Unified Entropy Control for Reinforcement Learning

SafetyDGX agent

arXiv:2604.14646v2 Announce Type: replace Abstract: Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (

The Amazing Stability of Flow Matching

ResearchDGX agent

arXiv:2604.16079v1 Announce Type: new Abstract: The success of deep generative models in generating high-quality and diverse samples is often attributed to particular architectures and large training

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world…

Model ReleasesDGX agent

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world trade, policy and controls work. And in some deeper sense h

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Model ReleasesDGX agent

arXiv:2604.16272v1 Announce Type: cross Abstract: As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured

When Cultures Meet: Multicultural Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2502.15972v2 Announce Type: replace-cross Abstract: Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultur

Where does output diversity collapse in post-training?

ResearchDGX agent

arXiv:2604.16027v1 Announce Type: cross Abstract: Post-trained language models produce less varied outputs than their base counterparts. This output diversity collapse undermines inference-time scalin

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

Model ReleasesDGX agent

arXiv:2510.17210v3 Announce Type: replace Abstract: The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Alon

18 Apr 2026

在 Ollama 运行 hermes

Local AiDGX agent

This post likely discusses how to run the Hermes language model using Ollama, an open-source tool for running large language models locally. It probably provides instructions or insights on setting up

17 Apr 2026

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

ResearchDGX agent

arXiv:2510.13829v3 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, reliable governance tools have become critical. Publicly verifiable watermarking is par

AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

Model ReleasesDGX agent

arXiv:2511.15915v2 Announce Type: replace-cross Abstract: We present AccelOpt, a self-improving large language model (LLM) agentic system that autonomously optimizes kernels for emerging AI acclerator

Anomaly Detection in IEC-61850 GOOSE Networks: Evaluating Unsupervised and Temporal Learning for Real-Time Intrusion Detection

ResearchDGX agent

arXiv:2604.14233v1 Announce Type: cross Abstract: The IEC-61850 GOOSE protocol underpins time-critical communication in modern digital substations but lacks native security mechanisms, leaving it vuln

Benchmarking Optimizers for MLPs in Tabular Deep Learning

Model ReleasesDGX agent

arXiv:2604.15297v1 Announce Type: new Abstract: MLP is a heavily used backbone in modern deep learning (DL) architectures for supervised learning on tabular data, and AdamW is the go-to optimizer used

CaptionQA: Is Your Caption as Useful as the Image Itself?

Model ReleasesDGX agent

arXiv:2511.21025v2 Announce Type: replace Abstract: Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic infe

Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate

Model ReleasesDGX agent

arXiv:2604.14210v1 Announce Type: new Abstract: A claim has been circulating on social media and practitioner forums that Chinese prompts are more token-efficient than English for LLM coding tasks, po

Claude Opus 4.7 on AI Gateway

Model ReleasesDGX agent

Claude Opus 4.7 is now available through Vercel's AI Gateway, Anthropic's latest large language model offering integration with Vercel's platform for developers. This enables developers to access Clau

Context Over Content: Exposing Evaluation Faking in Automated Judges

SafetyDGX agent

arXiv:2604.15224v1 Announce Type: cross Abstract: The extit{LLM-as-a-judge} paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: th

Data Synthesis Improves 3D Myotube Instance Segmentation

ResearchDGX agent

arXiv:2604.14720v1 Announce Type: new Abstract: Myotubes are multinucleated muscle fibers serving as key model systems for studying muscle physiology, disease mechanisms, and drug responses. Mechanist

Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance

ResearchDGX agent

arXiv:2604.14325v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance and have revolutionized NLP, but their lack of explainability keeps them treated as black boxes,

GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis

Model ReleasesDGX agent

arXiv:2604.13888v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into Geographic Information Systems (GIS) marks a paradigm shift toward autonomous spatial analysis. How

I have found 4.7 great for design, reverted back to 4.6 extended for everything else Anyone else like this?

Model ReleasesDGX agent

I have found 4.7 great for design, reverted back to 4.6 extended for everything else Anyone else like this? Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talk

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

Model ReleasesDGX agent

arXiv:2604.14799v1 Announce Type: new Abstract: Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing e

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

Model ReleasesDGX agent

arXiv:2604.15203v1 Announce Type: new Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Local AiDGX agent

arXiv:2505.18129v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodolog

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Model ReleasesDGX agent

arXiv:2506.03610v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game b

PeerPrism: Peer Evaluation Expertise vs Review-writing AI

Model ReleasesDGX agent

arXiv:2604.14513v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, ex

PixelDiT: Pixel Diffusion Transformers for Image Generation

ResearchDGX agent

arXiv:2511.20645v2 Announce Type: replace Abstract: Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoe

Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options

SafetyDGX agent

arXiv:2604.14634v1 Announce Type: new Abstract: Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by s

SAGE Celer 2.6 Technical Card

Model ReleasesDGX agent

arXiv:2604.14168v1 Announce Type: new Abstract: We introduce SAGE Celer 2.6, the latest in our line of general-purpose Celer models from SAGEA. Celer 2.6 is available in 5B, 10B, and 27B parameter siz

Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects

ApplicationsDGX agent

arXiv:2510.07890v3 Announce Type: replace Abstract: Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data. However, dialects are pri

The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

Local AiDGX agent

arXiv:2604.14703v1 Announce Type: new Abstract: Although some existing image manipulation localization (IML) methods incorporate authenticity-related supervision, this information is typically utilize

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

SafetyDGX agent

arXiv:2603.18373v2 Announce Type: replace Abstract: When VLMs answer correctly, do they genuinely rely on visual information or exploit language shortcuts? We introduce the Tri-Layer Diagnostic Framew

VeruSAGE: A Study of Agent-Based Verification for Rust Systems

Model ReleasesDGX agent

arXiv:2512.18436v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive capability to understand and develop code. However, their capability to rigorously reason a

xFODE+: Explainable Type-2 Fuzzy Additive ODEs for Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2604.14880v1 Announce Type: new Abstract: Recent advances in Deep Learning (DL) have boosted data-driven System Identification (SysID), but reliable use requires Uncertainty Quantification (UQ)

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception

SafetyDGX agent

arXiv:2510.23853v3 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlook

16 Apr 2026

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

Model ReleasesDGX agent

arXiv:2211.16780v3 Announce Type: replace-cross Abstract: In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because pre

Anthropic launches Claude Opus 4.7 with coding, visual reasoning improvements

Model ReleasesDGX agent

Anthropic PBC today opened access to Claude Opus 4.7, the latest addition to its popular line of large language models. The company says that the LLM is significantly better than its predecessor at co

Best Ollama models/settings for an 8GB VPS (CPU only, ARM)? Running into memory & looping issues.

Local AiDGX agent

This Reddit thread discusses running Ollama on a resource-constrained 8GB CPU-only ARM VPS, addressing common challenges such as out-of-memory errors and model response looping. For purely CPU-only se

Bi-Predictability: A Real-Time Signal for Monitoring LLM Interaction Integrity

AgentsDGX agent

arXiv:2604.13061v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in high-stakes autonomous and interactive workflows, where reliability demands continuous, multi-

Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text Detection

Model ReleasesDGX agent

arXiv:2604.13692v1 Announce Type: new Abstract: As large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from huma

Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon auton…

Model ReleasesDGX agent

Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon autonomy, unlocking a class of deep investigation work we couldn'

Codex for (almost) everything

Model ReleasesDGX agent

OpenAI's Codex is a large language model trained on publicly available code from the internet that can understand and generate code in dozens of programming languages. It powers GitHub Copilot and can

Counterfactual Peptide Editing for Causal TCR--pMHC Binding Inference

Model ReleasesDGX agent

arXiv:2604.13256v1 Announce Type: new Abstract: Neural models for TCR-pMHC binding prediction are susceptible to shortcut learning: they exploit spurious correlations in training data -- such as pepti

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

ResearchDGX agent

arXiv:2604.13286v1 Announce Type: new Abstract: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to p

Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression

SafetyDGX agent

arXiv:2604.13410v1 Announce Type: cross Abstract: We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A

Exposia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Model ReleasesDGX agent

arXiv:2601.06536v2 Announce Type: replace Abstract: We present Exposia, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded com

FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

Model ReleasesDGX agent

arXiv:2509.16445v2 Announce Type: replace Abstract: Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-wo

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

Model ReleasesDGX agent

arXiv:2602.23636v3 Announce Type: replace Abstract: Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed

Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card

Model ReleasesDGX agent

arXiv:2604.13466v1 Announce Type: cross Abstract: The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your…

Model ReleasesDGX agent

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your chat template. http://huggingface.co/zai-org/GLM-5.1/blob/m

Here's Qwen 3.6-35B-A3B v.s. Claude Opus 4.7 for 'Generate an SVG of a flamingo riding a unicycle', in case you thought Qwen might be cheati…

Model ReleasesDGX agent

This post compares the performance of Qwen 3.6-35B-A3B and Claude Opus 4.7 models on a creative task of generating SVG code for a flamingo riding a unicycle, likely demonstrating differences in their

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

Model ReleasesDGX agent

arXiv:2604.13058v1 Announce Type: new Abstract: We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,46

KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs

Model ReleasesDGX agent

arXiv:2604.13226v1 Announce Type: new Abstract: Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: re

Language steering in latent space to mitigate unintended code-switching

Model ReleasesDGX agent

arXiv:2510.13849v3 Announce Type: replace Abstract: Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks.

LM Performance:Qwen3.6-35B-A3B outperforms the dense 27B-param Qwen3.5-27B on several key coding benchmarks and dramatically surpasses its d…

AgentsDGX agent

LM Performance:Qwen3.6-35B-A3B outperforms the dense 27B-param Qwen3.5-27B on several key coding benchmarks and dramatically surpasses its direct predecessor Qwen3.5-35B-A3B, especially on agentic cod

LTX distilled 1.1 is the new king!

Local AiDGX agent

A Reddit thread from r/StableDiffusion discussing the release and community reception of LTX-Video's distilled 1.1 model, developed by Lightricks. LTX-Video is described as the first DiT-based video g

Need help setting up ollama.

Local AiDGX agent

A Reddit thread from the r/ollama community where a user seeks assistance with the initial setup and configuration of Ollama, a tool for running large language models locally. The discussion likely co

Nucleus Image now supported in Ostris' AI-Toolkit.

Local AiDGX agent

Ostris' AI Toolkit is an all-in-one training suite for diffusion models , and Nucleus Image has been added to the list of supported models . The toolkit can be run as a GUI or CLI and is designed to b

← Previous
1…367368369370371…1050
Next →