Applied AI Research Portfolio Get In Touch
Applied AI Research Lab

Frontier AI research.
Production AI systems.

We build AI products on frontier models, train custom models where APIs fall short, and take both to production — while publishing research that advances the field.

100+ enterprise projects Published @ EACL 2026 #1 on ViDoRe & ArmBench NeurIPS 2026 submission
What we do

Two things, done seriously.

100+
Projects delivered
8+
Years in production AI
#1
ArmBench & ViDoRe benchmarks
2026
NeurIPS & EACL publications
What we build

Applied AI capabilities

From custom model training to full production deployments — across every AI modality.

Custom LLMs
Multilingual · Domain-tuned
VLMs
Document · Video · Scene
VLA Models
Vision-Language-Action
Voice Agents
Multilingual · Real-time
AI Agents
Autonomous · Multi-agent
Computer Vision
Industrial · Medical
Semantic Search
RAG · Embeddings
Analytics & ML
Forecasting · BI · Tabular
Full Applied AI overview →
Research meets production

Physical AI. In production and in print.

We train JEPA-based world models with long-horizon supervision, build and benchmark robotic failure detection (FailBench, FailDetect), and run meta-analysis across the Physical AI benchmark landscape (MetaBench Physical). Papers are published or under review; models and benchmarks are headed to open source.

The same expertise ships in client work: anomaly detection, factory monitoring, defect detection, and traffic intelligence — deployed at production scale.

See our Physical AI research →
World Models
Action-conditioned JEPA trained with long-horizon and image-goal supervision
Failure Detection
FailBench & FailDetect — evaluating and detecting robotic arm failures
Benchmark Science
MetaBench Physical — one comparable matrix across Physical AI benchmarks
Industrial Vision
Defect detection, factory monitoring and traffic intelligence in production
Open science

Research highlights

All research →
NeurIPS 2026 Under review

Intuitive Physics & Planning with JEPA

Improving intuitive physics understanding through long-horizon supervision, plus action-conditioned JEPA world models steered by image goals and text instructions.

FailBench · FailDetect Open-sourcing

Robotic Failure Detection

A pooled benchmark for robotic arm failure detection — evaluating specialist models against generalist VLMs — and FailDetect, a model that zooms, crops, and catches failures.

MetaBench Physical Leaderboard

Physical AI Benchmark Meta-Analysis

Every vendor reports on different benchmarks. MetaBench Physical builds one comparable matrix, prunes redundant benchmarks to an optimal subset, and ranks models on it.

Also from the lab: low-resource text embeddings, published at EACL 2026 · #1 held on ViDoRe visual document retrieval · Uzbek legal embeddings & RAG, from a client engagement.

Browse Metric-AI on Hugging Face
Selected work

From the portfolio

Full portfolio →
Voice AI 2025 · Major Bank

Real-time Multilingual Fraud Detection Agent

Full voice pipeline — ASR, LLM, TTS — with fine-tuned models for Armenian, English and Russian. Sub-300ms latency, 98.7% accuracy.

Physical AI 2025 · Fortune 500 Manufacturer

Visual Defect Detection on Production Lines

Industrial anomaly detection deployed across 12 production lines. 40% reduction in defect pass-through.

GenAI & RAG 2024 · Series C Legal Tech

AI Legal Assistant with Long-Context Reasoning

Multimodal RAG pipeline over complex legal documents. Fine-tuned for jurisdiction-specific reasoning.

Get in touch

Tell us what you're building.

We respond to every serious inquiry — usually within a day.

Or email us directly at info@metric.am