Loading...
Browse 4,151 models across providers, modalities, and use cases.
🌐 All Models
4,151 models · Page 110 of 116
Ultra-fast, cost-efficient text-to-speech with ~120ms latency and 15-language support
Get embeddings for image using siglip-large-patch16-384
High-quality image generation and editing with support for eight reference images
Temporary publish target for bgogo deployment validation.
🔥 SeedVR2: one-step video & image restoration with 7B and Adjustable Resolution
Alpha-only video matting for 5-10 second white-background videos using BackgroundMattingV2 MobileNetV2 TorchScript.
Take a flat graphic, remove text, and get structured text layers back for editing and recomposing
A 4B-parameter safety and content moderation model that classifies user prompts and assistant responses as Safe, Unsafe, or Controversial with fine-grained category labels and refusal detection. Supports 119 languages.
GeoCalib (ECCV 2024): Single-image camera calibration. Estimates focal length, FoV, distortion, roll and pitch from one image using a deep net + Levenberg-Marquardt optimizer. Works on both outdoor and indoor scenes.
A custom Flux LoRA model trained on painterly illustrated poster art inspired by Blade Runner 2049. The style features atmospheric cyberpunk cityscapes with dramatic scale — tiny silhouetted figures dwarfed by massive holographic projections and towering
New and improved version of Veo 3 Fast, with higher-fidelity video, context-aware audio and last frame support
Generate professional e-commerce product photos from a single image. Automatically removes background, creates realistic studio scenes, and adds natural shadows.
The highest fidelity image model from Black Forest Labs
Create a dotted waveform video from an audio file
4x face image upscaler trained on FFHQ dataset using DAT (Dual Aggregation Transformer) architecture. Optimized for portrait and face photos.
Generate and edit images with Alibaba's Wan 2.7
A faster variant of Seedance 2.0 for quicker video generation with multimodal inputs and native audio.
This is a version of Qwen 3.5 27B optimised by Pruna AI.
High-fidelity video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
Ultralytics YOLOv8s worldv2 Real-Time Open-Vocabulary Object Detection model with 12.7M parameters. Achieves 37.7 mAP50-95 on COCO dataset. Optimized for real-time inference
This is Platmoji 2, trained more to mimic emojis in an extremely similar way. (Realism in emojis go to Platmoji 1)
This LUT-based color filter is ideal for color grading user-generated AI videos or short videos shot on smartphones.
RobustVideoMatting on Replicate: input mp4 video, output black-and-white alpha-mask.mp4.
ML model that detects acne, dark circles, wrinkles and oily skin in one go.
Image to separate stems from a song, using demucs and spleeter
Generate full-length songs up to 3 minutes from text prompts or images with Lyria 3 Pro, Google's most capable music generation model
Monocular metric depth estimation for panoramic images
Google's cost-efficient video generation model with native audio, optimized for high-volume applications
Beta/RFC version of https://replicate.com/nelsonjchen/op-replay-clipper
Enables precise control of character actions and expressions from a reference image.