AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
1 May 2026

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Model ReleasesDGX agent

arXiv:2604.27393v1 Announce Type: new Abstract: Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming inter

Modeling Spatial Extremal Dependence of Precipitation Using Distributional Neural Networks

ResearchDGX agent

arXiv:2407.08668v3 Announce Type: replace-cross Abstract: In this work, we propose a simulation-based estimation approach using generative neural networks to determine dependencies of precipitation ma

Musk v. Altman week 1: Elon Musk says he was duped, warns AI could kill us all, and admits that xAI distills OpenAI’s models

ResearchDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

In the first week of the landmark trial between Elon Musk and OpenAI, Musk took the stand in a crisp black suit and tie and argued that OpenAI CEO Sam Altman and president Greg Brockman had deceived h

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

ResearchDGX agent

arXiv:2604.27637v1 Announce Type: new Abstract: Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from t

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs

Model ReleasesDGX agent

arXiv:2604.27401v1 Announce Type: new Abstract: Perturbation probing generates task-specific causal hypotheses for FFN neurons in large language models using two forward passes per prompt and no backp

PVeRA: Probabilistic Vector-Based Random Matrix Adaptation

Model ReleasesDGX agent

arXiv:2512.07703v2 Announce Type: replace Abstract: Large foundation models have emerged in the last years and are pushing performance boundaries for a variety of tasks. Training or even finetuning su

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

Model ReleasesDGX agent

arXiv:2601.14289v2 Announce Type: replace-cross Abstract: Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables

RuC: HDL-Agnostic Rule Completion Benchmark Generation

Model ReleasesDGX agent

arXiv:2604.27780v1 Announce Type: cross Abstract: Large Language Models (LLMs) have rapidly improved in performance across code-related tasks, making their integration into Register Transfer Level (RT

SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images

Model ReleasesDGX agent

arXiv:2604.28039v1 Announce Type: new Abstract: Spectra are a prevalent yet highly information-dense form of scientific imagery, presenting substantial challenges to multimodal large language models (

30 Apr 2026

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio

Model ReleasesDGX agent

arXiv:2409.06624v4 Announce Type: replace-cross Abstract: Large Language Models (LLM) often need to be Continual Pre-Trained (CPT) to obtain unfamiliar language skills or adapt to new domains. The hug

A Systematic Comparison of Prompting and Multi-Agent Methods for LLM-based Stance Detection

Model ReleasesDGX agent

arXiv:2604.26319v1 Announce Type: new Abstract: Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task

AWS Generative AI Model Agility Solution: A comprehensive guide to migrating LLMs for generative AI production

TutorialsDGX agent

In this post, we introduce a systematic framework for LLM migration or upgrade in generative AI production, encompassing essential tools, methodologies, and best practices. The framework facilitates t

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what ou…

Model ReleasesDGX agent

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us' > all the existin

Bootstrapping Sign Language Annotations with Sign Language Models

ResearchDGX agent

AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of d

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

ResearchDGX agent

arXiv:2604.26169v1 Announce Type: new Abstract: Treatment allocation under budget constraints is a central challenge in digital advertising: advertisers must decide which users to show ads to while sp

CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs

Model ReleasesDGX agent

arXiv:2604.26378v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-p

Emergent Coordination in Multi-Agent Language Models

Local AiDGX agent

arXiv:2510.05174v4 Announce Type: replace-cross Abstract: When are multi-agent LLM systems merely a collection of individual agents versus an integrated collective with higher-order structure? We intr

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Model ReleasesDGX agent

arXiv:2604.26866v1 Announce Type: new Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of

Thinking with Drafting: Optical Decompression via Logical Reconstruction

Model ReleasesDGX agent

arXiv:2602.11731v2 Announce Type: replace Abstract: Existing multimodal large language models have achieved high-fidelity visual perception and exploratory visual generation. However, a precision para

VulStyle: A Multi-Modal Pre-Training for Code Stylometry-Augmented Vulnerability Detection

Model ReleasesDGX agent

arXiv:2604.26313v1 Announce Type: cross Abstract: We present VulStyle, a multi-modal software vulnerability detection model that jointly encodes function-level source code, non-terminal Abstract Synta

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

SafetyDGX agent

arXiv:2510.17548v2 Announce Type: replace Abstract: Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity,

29 Apr 2026

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

Model ReleasesDGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

Model ReleasesDGX agent

arXiv:2602.11786v2 Announce Type: replace Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation acr

FARM: Enhancing Molecular Representations with Functional Group Awareness

Model ReleasesDGX agent

arXiv:2410.02082v4 Announce Type: replace Abstract: We introduce Functional Group-Aware Representations for Small Molecules (FARM), a novel foundation model designed to bridge the gap between SMILES,

ollama run ministral-3:3b throwing error

Local AiDGX agent

The Ministral-3:3b model requires Ollama 0.13.1, which is in pre-release , and users encountering errors when running it face various issues including memory allocation problems and GPU/CPU offloading

Read the DeepSeek V4 Pro quickstart https://docs.together.ai/docs/deepseek-v4-quickstart

Model ReleasesDGX agent

DeepSeek V4 Pro is a language model available through Together AI's platform, with official quickstart documentation provided to help users get started with the model. The quickstart guide likely cove

The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents

Model ReleasesDGX agent

arXiv:2604.25299v1 Announce Type: new Abstract: Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following ta

Toward Multimodal Conversational AI for Age-Related Macular Degeneration

ResearchDGX agent

arXiv:2604.25720v1 Announce Type: cross Abstract: Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

SafetyDGX agent

arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi

28 Apr 2026

50+ fully managed MCP servers now available for Google Cloud services

Model ReleasesDGX agent

At Google Cloud Next ‘26, we announced that more than 50 Google-managed Model Context Protocol (MCP) servers are generally available or in preview, with more on the way. Why it matters: To move beyond

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

SafetyDGX agent

arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards

Model ReleasesDGX agent

arXiv:2604.22840v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a

Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge

Model ReleasesDGX agent

arXiv:2604.23701v1 Announce Type: cross Abstract: Crop disease diagnosis from field photographs faces two recurring problems: models that score well on benchmarks frequently hallucinate species names,

AI Safety Training Can be Clinically Harmful

SafetyDGX agent

arXiv:2604.23445v1 Announce Type: cross Abstract: Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigo

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code

Model ReleasesDGX agent

arXiv:2604.23361v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance on a wide range of software engineering tasks, including code generation and analysi

C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs

Model ReleasesDGX agent

arXiv:2604.23061v1 Announce Type: cross Abstract: Large language models (LLMs) show promise for molecular optimization, but aligning them with selective and competing drug-design constraints remains c

Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs

Model ReleasesDGX agent

arXiv:2509.17314v3 Announce Type: replace-cross Abstract: Software increasingly relies on the emergent capabilities of Large Language Models (LLMs), from natural language understanding to program anal

Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs

SafetyDGX agent

arXiv:2604.23949v1 Announce Type: new Abstract: Medical and public health experts must make real-time resource decisions, such as expanding hospital bed capacity, based on projected hospitalization tr

Don't Make the LLM Read the Graph: Make the Graph Think

Model ReleasesDGX agent

arXiv:2604.23057v1 Announce Type: new Abstract: We investigate whether explicit belief graphs improve LLM performance in cooperative multi-agent reasoning. Through 3,000+ controlled trials across four

DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2510.15050v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have made rapid progress, yet their reasoning ability often lags behind strong text-only LLMs. Bridging thi

GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2604.24477v1 Announce Type: cross Abstract: The rapid integration of Large Language Models (LLMs) into Multi-Agent Systems (MAS) has significantly enhanced their collaborative problem-solving ca

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

SafetyDGX agent

arXiv:2506.04118v3 Announce Type: replace Abstract: We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft be

I'm so confused…

Model ReleasesDGX agent

I'm so confused… We're excited to partner with Google to offer Grounding With Exa inside of Gemini models! Using Exa's agent-first search, Gemini models can now access billions of websites, technical

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

Model ReleasesDGX agent

NVIDIA's Nemotron 3 Nano Omni is a lightweight multimodal AI model capable of processing documents, audio, and video inputs for building intelligent agents. The model supports long-context understandi

Knowledge Vector of Logical Reasoning in Large Language Models

ResearchDGX agent

arXiv:2604.23877v1 Announce Type: new Abstract: Logical reasoning serve as a central capability in LLMs and includes three main forms: deductive, inductive, and abductive reasoning. In this work, we s

LILogic Net: Compact Logic Gate Networks with Learnable Connectivity for Efficient Hardware Deployment

ResearchDGX agent

arXiv:2511.12340v2 Announce Type: replace Abstract: Efficient machine learning deployment requires models that account for hardware constraints. Because binary logic gates are the fundamental primitiv

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling

Model ReleasesDGX agent

arXiv:2604.24715v1 Announce Type: new Abstract: Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transforme

MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

Model ReleasesDGX agent

arXiv:2511.14967v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Merma

Meta-CoT: Enhancing Granularity and Generalization in Image Editing

Model ReleasesDGX agent

arXiv:2604.24625v1 Announce Type: cross Abstract: Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into the

ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services

Model ReleasesDGX agent

arXiv:2604.24023v1 Announce Type: new Abstract: Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their p

Sphere-Depth: A Benchmark for Depth Estimation Methods with Varying Spherical Camera Orientations

Model ReleasesDGX agent

arXiv:2604.23432v1 Announce Type: cross Abstract: Reliable depth estimation from spherical images is crucial for 360{eg} vision in robotic navigation and immersive scene understanding. However, the on

The Pragmatic Persona: Discovering LLM Persona through Bridging Inference

Model ReleasesDGX agent

arXiv:2604.24079v1 Announce Type: cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue. However, most existing persona discovery approaches rely on su

Toward Theoretical Insights into Diffusion Trajectory Distillation via Operator Merging

ResearchDGX agent

arXiv:2505.16024v2 Announce Type: replace-cross Abstract: Diffusion trajectory distillation accelerates sampling by training a student model to approximate the multi-step denoising trajectories of a p

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks

Model ReleasesDGX agent

arXiv:2604.23145v1 Announce Type: cross Abstract: Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent comple

YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review

ResearchDGX agent

arXiv:2501.13400v3 Announce Type: replace-cross Abstract: In the field of deep learning-based computer vision, YOLO is revolutionary. With respect to deep learning models, YOLO is also the one that is

Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data

Model ReleasesDGX agent

arXiv:2604.24479v1 Announce Type: new Abstract: Computer-Aided Design (CAD) models are defined by their construction history: a parametric recipe that encodes design intent. However, existing large-sc

27 Apr 2026

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution

Model ReleasesDGX agent

arXiv:2604.22192v1 Announce Type: new Abstract: Chart-to-code generation demands strict visual precision and syntactic correctness from Vision-Language Models (VLMs). However, existing approaches are

CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

Model ReleasesDGX agent

arXiv:2604.22367v1 Announce Type: cross Abstract: Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs t

Conditional Diffusion Posterior Alignment for Sparse-View CT Reconstruction

SafetyDGX agent

arXiv:2604.21960v1 Announce Type: cross Abstract: Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, t

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

SafetyDGX agent

arXiv:2604.22119v1 Announce Type: new Abstract: As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own ob

← Previous
1…289290291292293…1042
Next →