- 
	
	
	
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 28 - 
	
	
	
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 14 - 
	
	
	
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 44 - 
	
	
	
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 23 
Collections
Discover the best community collections!
Collections including paper arxiv:2411.17949 
						
					
				- 
	
	
	266
Hunyuan3D-1.0
😻Text-to-3D and Image-to-3D Generation
 - 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
Paper • 2502.02492 • Published • 66 - 
	
	
	
Phantom: Subject-consistent video generation via cross-modal alignment
Paper • 2502.11079 • Published • 59 
- 
	
	
	
LinFusion: 1 GPU, 1 Minute, 16K Image
Paper • 2409.02097 • Published • 34 - 
	
	
	
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
Paper • 2409.11406 • Published • 27 - 
	
	
	
Diffusion Models Are Real-Time Game Engines
Paper • 2408.14837 • Published • 126 - 
	
	
	
Segment Anything with Multiple Modalities
Paper • 2408.09085 • Published • 22 
- 
	
	
	
Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis
Paper • 2401.09048 • Published • 10 - 
	
	
	
Improving fine-grained understanding in image-text pre-training
Paper • 2401.09865 • Published • 18 - 
	
	
	
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Paper • 2401.10891 • Published • 62 - 
	
	
	
Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
Paper • 2401.13627 • Published • 77 
- 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
Unpacking SDXL Turbo: Interpreting Text-to-Image Models with Sparse Autoencoders
Paper • 2410.22366 • Published • 83 - 
	
	
	
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Paper • 2410.13863 • Published • 38 - 
	
	
	
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
Paper • 2408.15914 • Published • 24 
- 
	
	
	
Generating Compositional Scenes via Text-to-image RGBA Instance Generation
Paper • 2411.10913 • Published • 4 - 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
Pathways on the Image Manifold: Image Editing via Video Generation
Paper • 2411.16819 • Published • 37 
- 
	
	
	
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Paper • 2405.08748 • Published • 24 - 
	
	
	
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper • 2405.10300 • Published • 30 - 
	
	
	
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Paper • 2405.09818 • Published • 131 - 
	
	
	
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Paper • 2405.11143 • Published • 41 
- 
	
	
	
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 28 - 
	
	
	
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 14 - 
	
	
	
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 44 - 
	
	
	
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 23 
- 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
Unpacking SDXL Turbo: Interpreting Text-to-Image Models with Sparse Autoencoders
Paper • 2410.22366 • Published • 83 - 
	
	
	
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Paper • 2410.13863 • Published • 38 - 
	
	
	
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
Paper • 2408.15914 • Published • 24 
- 
	
	
	266
Hunyuan3D-1.0
😻Text-to-3D and Image-to-3D Generation
 - 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
Paper • 2502.02492 • Published • 66 - 
	
	
	
Phantom: Subject-consistent video generation via cross-modal alignment
Paper • 2502.11079 • Published • 59 
- 
	
	
	
Generating Compositional Scenes via Text-to-image RGBA Instance Generation
Paper • 2411.10913 • Published • 4 - 
	
	
	
ROICtrl: Boosting Instance Control for Visual Generation
Paper • 2411.17949 • Published • 87 - 
	
	
	
Pathways on the Image Manifold: Image Editing via Video Generation
Paper • 2411.16819 • Published • 37 
- 
	
	
	
LinFusion: 1 GPU, 1 Minute, 16K Image
Paper • 2409.02097 • Published • 34 - 
	
	
	
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
Paper • 2409.11406 • Published • 27 - 
	
	
	
Diffusion Models Are Real-Time Game Engines
Paper • 2408.14837 • Published • 126 - 
	
	
	
Segment Anything with Multiple Modalities
Paper • 2408.09085 • Published • 22 
- 
	
	
	
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Paper • 2405.08748 • Published • 24 - 
	
	
	
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper • 2405.10300 • Published • 30 - 
	
	
	
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Paper • 2405.09818 • Published • 131 - 
	
	
	
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Paper • 2405.11143 • Published • 41 
- 
	
	
	
Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis
Paper • 2401.09048 • Published • 10 - 
	
	
	
Improving fine-grained understanding in image-text pre-training
Paper • 2401.09865 • Published • 18 - 
	
	
	
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Paper • 2401.10891 • Published • 62 - 
	
	
	
Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
Paper • 2401.13627 • Published • 77