AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
7 Aug 2026

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

Model ReleasesDGX agent

arXiv:2608.05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Model ReleasesDGX agent

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Model ReleasesDGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-trainin

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

SafetyDGX agent

arXiv:2608.06125v1 Announce Type: new Abstract: Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human prefer

SEAM: Global consistency beyond local accuracy in scientific machine learning

Model ReleasesDGX agent

arXiv:2608.05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. Yet such loc

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

Model ReleasesDGX agent

arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining rema

Timestep-Conditioned Transformers for Global Weather Forecasting

ResearchDGX agent

arXiv:2608.06241v1 Announce Type: new Abstract: Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a f

Tree-NET: Enhancing 2D Medical Image Segmentation Through Efficient Low-Level Feature Training

Model ReleasesDGX agent

arXiv:2501.02140v2 Announce Type: replace-cross Abstract: This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segment

6 Aug 2026

Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

Model ReleasesDGX agent

arXiv:2608.04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual

critical history and context that a lot of people have conveniently forgotten

SafetyDGX agent

critical history and context that a lot of people have conveniently forgotten For a very long time most high-performing AI models were end-to-end neural models; vector input -> vector output, with onl

Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

Model ReleasesDGX agent

arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high

Foreseeing the Invisible: Amodal Reconstruction of Leaf Fossil Images

Model ReleasesDGX agent

arXiv:2608.04423v1 Announce Type: new Abstract: Fossil leaves are rarely preserved whole -- sedimentary rock hides, breaks, and erodes the lamina, yet paleobotany depends on the complete shape and out

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Model ReleasesDGX agent

arXiv:2608.05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tunin

How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode

Model ReleasesDGX agent

Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

Model ReleasesDGX agent

arXiv:2608.04519v1 Announce Type: new Abstract: Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Curre

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

SafetyDGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding

Model ReleasesDGX agent

arXiv:2604.00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Alth

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

Model ReleasesDGX agent

arXiv:2608.04791v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization

Outlook where frontier AI is headed next 18 months: The AI reasoning training + harness loop works if you can produce enough data and reason…

AgentsDGX agent

Outlook where frontier AI is headed next 18 months: The AI reasoning training + harness loop works if you can produce enough data and reasoning traces (via verifiers). Proven with code and math result

Protoreasoning in Tiny Transformers

Model ReleasesDGX agent

arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-ste

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Model ReleasesDGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

Right Reset: Chunking by Prefix Removal

ResearchDGX agent

arXiv:2608.04330v1 Announce Type: new Abstract: Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens w

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Model ReleasesDGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals

Model ReleasesDGX agent

arXiv:2608.04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable

Thinking with Anchors: Grounded and Efficient Document Reasoning

Model ReleasesDGX agent

arXiv:2608.04424v1 Announce Type: new Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reaso

5 Aug 2026

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

Model ReleasesDGX agent

arXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in eval

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

AgentsDGX agent

arXiv:2608.00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominan

Approximate Speculative Decoding

Model ReleasesDGX agent

arXiv:2608.03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verificat

Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling

ResearchDGX agent

arXiv:2608.02618v1 Announce Type: new Abstract: Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized con

CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

SafetyDGX agent

arXiv:2608.02663v1 Announce Type: cross Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

Model ReleasesDGX agent

arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to

Disentangling MLP Neuron Weights in Vocabulary Space

Model ReleasesDGX agent

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks

Model ReleasesDGX agent

arXiv:2608.03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant infor

Don't Walk the Line: Boundary Guidance for Filtered Generation

Model ReleasesDGX agent

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tun

Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems

ApplicationsDGX agent

arXiv:2608.03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image gen

Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

Model ReleasesDGX agent

arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administra

Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors

Model ReleasesDGX agent

arXiv:2511.15968v2 Announce Type: replace-cross Abstract: External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shifts

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

Model ReleasesDGX agent

arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks rem

FLARE: Few-shot Learning-based Adaptive Reflective Engine

Model ReleasesDGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Model ReleasesDGX agent

arXiv:2608.03826v1 Announce Type: new Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations,

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

Model ReleasesDGX agent

arXiv:2608.03782v1 Announce Type: new Abstract: Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus o

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

Model ReleasesDGX agent

arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review,

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews

Model ReleasesDGX agent

arXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smo

LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics

ApplicationsDGX agent

arXiv:2603.24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation

PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning

Model ReleasesDGX agent

arXiv:2608.03034v1 Announce Type: cross Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains imp

Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation

Model ReleasesDGX agent

arXiv:2608.03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

Model ReleasesDGX agent

arXiv:2608.02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying mode

Rubrics as Privileged Information for Open-Ended Generation

Model ReleasesDGX agent

arXiv:2608.02948v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable dom

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

Model ReleasesDGX agent

arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goe

Separating quantum circuits from classical LLMs

ResearchDGX agent

arXiv:2608.03962v1 Announce Type: cross Abstract: Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generatio

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

TutorialsDGX agent

arXiv:2608.03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Model ReleasesDGX agent

arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

Model ReleasesDGX agent

arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research fro

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

Model ReleasesDGX agent

arXiv:2608.03084v1 Announce Type: new Abstract: End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output form

Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection

Model ReleasesDGX agent

arXiv:2608.03627v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexp

4 Aug 2026

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2509.24560v2 Announce Type: replace Abstract: Extended Chain-of-Thought (CoT) reasoning has significantly bolstered the capabilities of medical large language models (LLMs). However, current mod

ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces

Model ReleasesDGX agent

arXiv:2608.01968v1 Announce Type: new Abstract: Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. Yet

Control Under Compression: Reliability Frontiers for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

Model ReleasesDGX agent

arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation

← Previous
1…296297298299300…1036
Next →