AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,429Total entries
1Added by human
88,428Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,648 results
Model Releases

Dynamic Scaled Gradient Descent for Stable Fine-Tuning for Classifications

DGX agent

arXiv:2604.27987v1 Announce Type: new Abstract: Fine-tuning pretrained models has become a standard approach to adapting pretrained knowledge to improve the accuracy on new sparse, imbalance datasets.

model-releasesarxiv-cs-lg
1 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Exploration Hacking: Can LLMs Learn to Resist RL Training?

DGX agent

arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignmen

safetyarxiv-cs-cl
1 May 2026
Model Releases

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs

DGX agent

arXiv:2506.07180v3 Announce Type: replace-cross Abstract: As video large language models (Video-LLMs) become increasingly integrated into real-world applications that demand grounded multimodal reason

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation

DGX agent

arXiv:2604.27969v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate visual artifacts into code, from UI mockups into HTML to scientific plots

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

DGX agent

arXiv:2604.27906v1 Announce Type: new Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant contex

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation

DGX agent

arXiv:2604.27249v1 Announce Type: cross Abstract: When instructed to underperform on multiple-choice evaluations, do language models engage with question content or fall back on positional shortcuts?

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets

DGX agent

arXiv:2509.15549v2 Announce Type: replace Abstract: Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, hi

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment

DGX agent

arXiv:2506.22500v2 Announce Type: replace-cross Abstract: Automated identification of surgical safety risks is critical for improving patient outcomes; however, Multimodal Large Language Models (MLLMs

model-releasesarxiv-cs-ai
1 May 2026
Safety

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor

DGX agent

arXiv:2604.27633v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated for political bias based on their responses to fixed questionnaires, which typically place frontier

safetyarxiv-cs-ai
1 May 2026
Model Releases

PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning

DGX agent

arXiv:2604.28123v1 Announce Type: cross Abstract: The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinfo

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner

DGX agent

arXiv:2509.09513v2 Announce Type: replace-cross Abstract: Biophysical diffusion MRI models like Neurite Exchange Imaging (NEXI) are essential for probing gray matter microstructure, estimating compart

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching

DGX agent

arXiv:2604.27375v1 Announce Type: new Abstract: Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise ret

model-releasesarxiv-cs-cv
1 May 2026
Model Releases

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

DGX agent

arXiv:2604.27228v1 Announce Type: new Abstract: Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles t

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Allen AI just released the OlmPool research series on Hugging Face Early 7-8B checkpoints trained to 150B tokens exploring how minor archite…

DGX agent

Allen AI released the OlmPool research series on Hugging Face, featuring early 7-8B parameter language model checkpoints trained on 150 billion tokens. The research explores how minor architectural mo

model-releasesclem-delangue--x
30 Apr 2026
Model Releases

COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data

DGX agent

arXiv:2603.03239v2 Announce Type: replace Abstract: Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land-cover. Relationships b

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces

DGX agent

arXiv:2505.18441v2 Announce Type: replace Abstract: Dictionary learning has recently emerged as a promising approach for mechanistic interpretability of large transformer models. Disentangling high-di

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

DSIPA: Detecting LLM-Generated Texts via Sentiment-Invariant Patterns Divergence Analysis

DGX agent

arXiv:2604.26328v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) presents new security challenges, particularly in detecting machine-generated text used for misi

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Efficient, VRAM-Constrained xLM Inference on Clients

DGX agent

arXiv:2604.26334v1 Announce Type: cross Abstract: To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language mo

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

DGX agent

arXiv:2602.03558v2 Announce Type: replace-cross Abstract: Generative text-to-image models are advancing at an unprecedented pace, continuously shifting the perceptual quality ceiling and rendering pre

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents

DGX agent

arXiv:2604.26274v1 Announce Type: cross Abstract: Structured-workflow agents driven by large language models execute tool calls against sensitive external environments. We propose odename, a telemetry

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

DGX agent

arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requireme

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing

DGX agent

arXiv:2601.21459v4 Announce Type: replace-cross Abstract: LLM role-playing, i.e., using LLMs to simulate specific personas, has emerged as a key capability in various applications, such as companionsh

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

MARVIS: Modality Adaptive Reasoning over VISualizations

DGX agent

arXiv:2507.01544v2 Announce Type: replace Abstract: Predictive applications of machine learning often rely on small (sub 1 Bn parameter) specialized models tuned to particular domains or modalities. S

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

DGX agent

arXiv:2601.02731v3 Announce Type: replace-cross Abstract: Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers signifi

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Perception Test 2025: Challenge Summary and a Unified VQA Extension

DGX agent

arXiv:2601.06287v2 Announce Type: replace Abstract: The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

DGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

DGX agent

arXiv:2604.26102v1 Announce Type: cross Abstract: Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

This startup’s new mechanistic interpretability tool lets you debug LLMs

DGX agent

The San Francisco–based startup Goodfire just released a new tool, called Silico, that lets researchers and engineers peer inside an AI model and adjust its parameters—the settings that determine a mo

model-releasesmit-tech-review
30 Apr 2026
Model Releases

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation

DGX agent

arXiv:2602.11224v3 Announce Type: replace-cross Abstract: We present Agent-Diff, a novel benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world productivity software

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

DGX agent

arXiv:2602.23069v2 Announce Type: replace Abstract: Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are fa

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Analyzing LLM Reasoning to Uncover Mental Health Stigma

DGX agent

arXiv:2604.25053v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma to

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering

DGX agent

arXiv:2601.12248v2 Announce Type: replace-cross Abstract: Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks m

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Architecture Determines Observability in Transformers

DGX agent

arXiv:2604.24801v1 Announce Type: new Abstract: Autoregressive transformers make confident errors, but activation monitoring can catch them only if the model preserves an internal signal that output c

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Cross-Lingual Jailbreak Detection via Semantic Codebooks

DGX agent

arXiv:2604.25716v1 Announce Type: new Abstract: Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment.

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot

DGX agent

arXiv:2601.02078v2 Announce Type: replace Abstract: The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training da

model-releasesarxiv-cs-ro
29 Apr 2026
Model Releases

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

DGX agent

arXiv:2604.25361v1 Announce Type: new Abstract: Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluati

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Investigation into In-Context Learning Capabilities of Transformers

DGX agent

arXiv:2604.25858v1 Announce Type: new Abstract: Transformers have demonstrated a strong ability for in-context learning (ICL), enabling models to solve previously unseen tasks using only example input

model-releasesarxiv-cs-lg
29 Apr 2026
Applications

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

DGX agent

arXiv:2603.29080v2 Announce Type: replace Abstract: Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existi

applicationsarxiv-cs-cv
29 Apr 2026
Local Ai

Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling

DGX agent

arXiv:2604.25860v1 Announce Type: new Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specif

local-aiarxiv-cs-cl
29 Apr 2026
Model Releases

M^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

DGX agent

arXiv:2604.25122v1 Announce Type: new Abstract: We present M^3-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (ML

model-releasesarxiv-cs-cv
29 Apr 2026
Research

Magnification-Invariant Image Classification via Domain Generalization and Stable Sparse Embedding Signatures

DGX agent

arXiv:2604.25817v1 Announce Type: new Abstract: Magnification shift is a major obstacle to robust histopathology classification, because models trained on one imaging scale often generalize poorly to

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Mitigating Coordinate Prediction Bias from Positional Encoding Failures

DGX agent

arXiv:2510.22102v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, precise coordinate prediction remains a significant cha

model-releasesarxiv-cs-cl
29 Apr 2026
Local Ai

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

DGX agent

arXiv:2603.15954v2 Announce Type: replace Abstract: Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The

local-aiarxiv-cs-lg
29 Apr 2026
Model Releases

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

DGX agent

arXiv:2604.24954v1 Announce Type: cross Abstract: We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, i

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

DGX agent

arXiv:2604.24964v1 Announce Type: cross Abstract: Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real wor

model-releasesarxiv-cs-cl
29 Apr 2026
Research

PLMGH: What Matters in PLM-GNN Hybrids for Code Classification and Vulnerability Detection

DGX agent

arXiv:2604.25599v1 Announce Type: cross Abstract: Code understanding models increasingly rely on pretrained language models (PLMs) and graph neural networks (GNNs), which capture complementary semanti

researcharxiv-cs-lg
29 Apr 2026
Tools

Qwen3.6-Plus is now available on Together AI Try it now: http://www.together.ai/models/qwen36-plus

DGX agent

Qwen3.6-Plus, a large language model, is now available for use through Together AI's platform. Together AI has announced the availability of this model and is inviting users to try it via their models

toolstogether-ai--x
29 Apr 2026
Model Releases

RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation

DGX agent

arXiv:2603.09723v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used across the scientific workflow, including to draft peer-review reports. However, many AI-generate

model-releasesarxiv-cs-cl
29 Apr 2026
← Previous
1…405406407408409…1326
Next →