AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,272 results
14 Apr 2026

LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-Tuning

Model ReleasesDGX agent

arXiv:2502.14644v5 Announce Type: replace Abstract: Long context understanding remains challenging for large language models due to their limited context windows. This paper introduces Long Input Fine

LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2511.14774v3 Announce Type: replace-cross Abstract: Evaluating cross-lingual knowledge transfer in large language models is challenging, as correct answers in a target language may arise either

llama.cpp built from source + qwen3.5 27b running locally + hermes agent on top + camofox for scraping (tls spoofing) + scweet to scrape x (…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

llama.cpp built from source + qwen3.5 27b running locally + hermes agent on top + camofox for scraping (tls spoofing) + scweet to scrape x (no api keys) + tailscale to access from other devices bro Me

LLM Knowledge Base → Slides When @karpathy shared his LLM Knowledge Base setup, many were wondering how to generate more visual forms of the…

Model ReleasesDGX agent

LLM Knowledge Base → Slides When @karpathy shared his LLM Knowledge Base setup, many were wondering how to generate more visual forms of the wiki. There are many options, but I think @GammaApp is one

LLMs for Text-Based Exploration and Navigation Under Partial Observability

Model ReleasesDGX agent

arXiv:2604.09604v1 Announce Type: new Abstract: Exploration and goal-directed navigation in unknown layouts are central to inspection, logistics, and search-and-rescue. We ask whether large language m

LLMs Should Incorporate Explicit Mechanisms for Human Empathy

Model ReleasesDGX agent

arXiv:2604.10557v1 Announce Type: cross Abstract: This paper argues that Large Language Models (LLMs) should incorporate explicit mechanisms for human empathy. As LLMs become increasingly deployed in

Local tool for cli coding like Claude code

Model ReleasesDGX agent

This r/ollama thread discusses how to run a local, free alternative to Claude Code for CLI-based AI coding using Ollama. Ollama v0.14.0 and later are compatible with the Anthropic Messages API, making

LoGo-MR: Screening Breast MRI for Cancer Risk Prediction by Efficient Omni-Slice Modeling

Model ReleasesDGX agent

arXiv:2604.11348v1 Announce Type: new Abstract: Efficient and explainable breast cancer (BC) risk prediction is critical for large-scale population-based screening. Breast MRI provides functional info

LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval

Model ReleasesDGX agent

arXiv:2601.14706v3 Announce Type: replace Abstract: In this paper, we present LookBench (We use the term 'look' to reflect retrieval that mirrors how people shop -- finding the exact item, a close sub

LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention

Model ReleasesDGX agent

arXiv:2604.10044v1 Announce Type: new Abstract: Through systematic experiments on long-context generation, we observe a damaging failure mode in which decoding can collapse into persistent repetition

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

Model ReleasesDGX agent

arXiv:2604.11792v1 Announce Type: new Abstract: Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multim

LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment: Methods and Results

Model ReleasesDGX agent

arXiv:2604.11207v1 Announce Type: new Abstract: This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration

Model ReleasesDGX agent

arXiv:2604.11446v1 Announce Type: cross Abstract: Recently, scaling reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs) has emerged as an effective training paradigm

LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection

Model ReleasesDGX agent

arXiv:2604.10862v1 Announce Type: new Abstract: The rapid advancement of diffusion-based generative models has made face forgery detection a critical challenge in digital forensics. Current detection

LumiMotion: Improving Gaussian Relighting with Scene Dynamics

Model ReleasesDGX agent

arXiv:2604.10994v1 Announce Type: new Abstract: In 3D reconstruction, the problem of inverse rendering, namely recovering the illumination of the scene and the material properties, is fundamental. Exi

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

Model ReleasesDGX agent

arXiv:2604.10024v1 Announce Type: cross Abstract: Long video summarization presents significant challenges for current multimodal large language models (MLLMs), particularly in maintaining temporal fi

M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency

Model ReleasesDGX agent

arXiv:2604.01306v2 Announce Type: replace Abstract: Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing

M2.7 w/ hermes cli is replacing ~75% of my claude code / opus usage now, but we need clarity for using it as a coding agent @ work. We're tr…

Model ReleasesDGX agent

M2.7 w/ hermes cli is replacing ~75% of my claude code / opus usage now, but we need clarity for using it as a coding agent @ work. We're truly blessed to have the weights of this one, looking forward

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

Model ReleasesDGX agent

arXiv:2511.18373v2 Announce Type: replace Abstract: Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial

MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis

Model ReleasesDGX agent

arXiv:2604.11188v1 Announce Type: cross Abstract: Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed

MAVEN-T: Multi-Agent enVironment-aware Enhanced Neural Trajectory predictor with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.10169v1 Announce Type: new Abstract: Trajectory prediction remains a critical yet challenging component in autonomous driving systems, requiring sophisticated reasoning capabilities while m

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages

Model ReleasesDGX agent

arXiv:2512.01512v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved great success in Speech-to-Text Translation (S2TT) tasks. However, current research is constr

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

Model ReleasesDGX agent

arXiv:2604.09552v1 Announce Type: cross Abstract: Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retri

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness

Model ReleasesDGX agent

arXiv:2603.22816v3 Announce Type: replace-cross Abstract: Language models increasingly show their work by writing step-by-step reasoning before answering. But are these steps genuinely used, or is the

Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models

Model ReleasesDGX agent

arXiv:2604.11216v1 Announce Type: new Abstract: What values, evidence preferences, and source trust hierarchies do AI systems actually exhibit when facing structured dilemmas? We present the first lar

Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation

Model ReleasesDGX agent

arXiv:2604.05795v2 Announce Type: replace Abstract: The increasing use of large language models in mental health applications calls for principled evaluation frameworks that assess alignment with psyc

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2602.21950v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown great potential in medical applications, yet existing benchmarks inadequately capture real-world

MedVeriSeg: Teaching MLLM-Based Medical Segmentation Models to Verify Query Validity Without Extra Training

Model ReleasesDGX agent

arXiv:2604.10242v1 Announce Type: new Abstract: Despite recent advances in MLLM-based medical image segmentation, existing LISA-like methods cannot reliably reject false queries and often produce hall

MemDLM: Memory-Enhanced DLM Training

Model ReleasesDGX agent

arXiv:2603.22241v2 Announce Type: replace Abstract: Diffusion Language Models (DLMs) offer attractive advantages over Auto-Regressive (AR) models, such as full-attention parallel decoding and flexible

MEMENTO: Teaching LLMs to Manage Their Own Context

Model ReleasesDGX agent

arXiv:2604.09852v1 Announce Type: new Abstract: Reasoning models think in long, unstructured streams with no mechanism for compressing or organizing their own intermediate state. We introduce MEMENTO:

MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment

Model ReleasesDGX agent

arXiv:2601.22361v2 Announce Type: replace-cross Abstract: Assessing the veracity of online content has become increasingly critical. Large language models (LLMs) have recently enabled substantial prog

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2604.11502v1 Announce Type: cross Abstract: Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate th

Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics

Model ReleasesDGX agent

arXiv:2604.11012v1 Announce Type: new Abstract: The quality of text generated by large language models depends critically on the decoding sampling strategy. While mainstream methods such as Top-k, Top

Minimizing classical resources in variational measurement-based quantum computation for generative modeling

Model ReleasesDGX agent

arXiv:2604.11578v1 Announce Type: cross Abstract: Measurement-based quantum computation (MBQC) is a framework for quantum information processing in which a computational task is carried out through on

Mintlify, which uses AI to help companies generate software documentation, raised a 45M Series B led by a16z and Salesforce Ventures at a 500M valuation (Rashi Shrivastava/Forbes)

Model ReleasesDGX agent

Rashi Shrivastava / Forbes: Mintlify, which uses AI to help companies generate software documentation, raised a 45M Series B led by a16z and Salesforce Ventures at a 500M valuation — It's just one of

Mirai: Autoregressive Visual Generation Needs Foresight

Model ReleasesDGX agent

arXiv:2601.14671v2 Announce Type: replace Abstract: Autoregressive (AR) visual generators model images as sequences of discrete tokens and are trained with a next-token likelihood objective. This stri

MLLM-as-a-Judge Exhibits Model Preference Bias

Model ReleasesDGX agent

arXiv:2604.11589v1 Announce Type: new Abstract: Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model perf

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.10971v1 Announce Type: cross Abstract: In the progress of industrial anomaly detection, general anomaly detection (GAD) is an emerging trend and also the ultimate goal. Unlike the conventio

MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

Model ReleasesDGX agent

arXiv:2604.10755v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely unte

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization

Model ReleasesDGX agent

arXiv:2604.11259v1 Announce Type: new Abstract: Mobile GUI agents powered by Multimodal Large Language Models (MLLMs) can execute complex tasks on mobile devices. Despite this progress, most existing

MoEITS: A Green AI approach for simplifying MoE-LLMs

Model ReleasesDGX agent

arXiv:2604.10603v1 Announce Type: cross Abstract: Large language models are transforming all areas of academia and industry, attracting the attention of researchers, professionals, and the general pub

MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI

Model ReleasesDGX agent

arXiv:2604.11762v1 Announce Type: new Abstract: Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been dr

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a s…

Model ReleasesDGX agent

Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real t

MPAC: A Multi-Principal Agent Coordination Protocol for Interoperable Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2604.09744v1 Announce Type: cross Abstract: The AI agent ecosystem has converged on two protocols: the Model Context Protocol (MCP) for tool invocation and Agent-to-Agent (A2A) for single-princi

MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling Analysis

Model ReleasesDGX agent

arXiv:2604.10126v1 Announce Type: cross Abstract: Metamorphic testing (MT) is a widely recognized technique for alleviating the oracle problem in software testing. However, its adoption is hindered by

Multi-Head Attention based interaction-aware architecture for Bangla Handwritten Character Recognition: Introducing a Primary Dataset

Model ReleasesDGX agent

arXiv:2604.09717v1 Announce Type: new Abstract: Character recognition is the fundamental part of an optical character recognition (OCR) system. Word recognition, sentence transcription, document digit

Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi

Model ReleasesDGX agent

arXiv:2604.10347v1 Announce Type: new Abstract: Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creatin

Multi-ORFT: Stable Online Reinforcement Fine-Tuning for Multi-Agent Diffusion Planning in Cooperative Driving

Model ReleasesDGX agent

arXiv:2604.11734v1 Announce Type: cross Abstract: Closed-loop cooperative driving requires planners that generate realistic multimodal multi-agent trajectories while improving safety and traffic effic

Muon^2: Boosting Muon via Adaptive Second-Moment Preconditioning

Model ReleasesDGX agent

arXiv:2604.09967v1 Announce Type: cross Abstract: Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates t

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

Model ReleasesDGX agent

arXiv:2512.00336v2 Announce Type: replace Abstract: The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content auth

Mycelium-Index: A Streaming Approximate Nearest Neighbor Index with Myelial Edge Decay, Traffic-Driven Reinforcement, and Adaptive Living Hierarchy

Model ReleasesDGX agent

arXiv:2604.11274v1 Announce Type: new Abstract: We present mycelium-index, a streaming approximate nearest neighbor (ANN) index for high-dimensional vector spaces, inspired by the adaptive growth patt

Nationality encoding in language model hidden states: Probing culturally differentiated representations in persona-conditioned academic text

Model ReleasesDGX agent

arXiv:2604.10151v1 Announce Type: new Abstract: Large language models are increasingly used as writing tools and pedagogical resources in English for Academic Purposes, but it remains unclear whether

Natural Gradient Gaussian Approximation Filter on Lie Groups for Robot State Estimation

Model ReleasesDGX agent

arXiv:2604.10057v1 Announce Type: new Abstract: Accurate state estimation for robotic systems evolving on Lie group manifolds, such as legged robots, is a prerequisite for achieving agile control. How

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration

Model ReleasesDGX agent

arXiv:2604.09678v1 Announce Type: cross Abstract: As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address

Neural Generalized Mixed-Effects Models

Model ReleasesDGX agent

arXiv:2604.10976v1 Announce Type: cross Abstract: Generalized linear mixed-effects models (GLMMs) are widely used to analyze grouped and hierarchical data. In a GLMM, each response is assumed to follo

Neural Stochastic Processes for Satellite Precipitation Refinement

Model ReleasesDGX agent

arXiv:2604.10414v1 Announce Type: new Abstract: Accurate precipitation estimation is critical for flood forecasting, water resource management, and disaster preparedness. Satellite products provide gl

NEW: Added @OpenRouter's new free stealth model, Elephant-Alpha, to Hermes Agent, which you can now access if you run `hermes update`! I als…

Model ReleasesDGX agent

NEW: Added @OpenRouter's new free stealth model, Elephant-Alpha, to Hermes Agent, which you can now access if you run `hermes update`! I also had Hermes come up with an agentic benchmark on the fly to

New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework

Model ReleasesDGX agent

arXiv:2604.09940v1 Announce Type: new Abstract: Fine-tuning Large Language Models (LLMs) typically involves either full fine-tuning, which updates all model parameters, or Parameter-Efficient Fine-Tun

New research result: we use Claude to make fully autonomous progress on scalable oversight research, as measured by performance gap recovere…

Model ReleasesDGX agent

New research result: we use Claude to make fully autonomous progress on scalable oversight research, as measured by performance gap recovered (PGR). Claude iterates on a number of different techniques

NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

Model ReleasesDGX agent

arXiv:2604.11543v1 Announce Type: cross Abstract: Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing p

← Previous
1…352353354355356…372
Next →