# Study: Video Datasets

Checking out some video datasets and evaluating them.

**Disclaimer**: This is not comprehensive evaluation. Just me manually looking through huggingface.

---

### PE Video

Link: [facebook/PE-Video](https://huggingface.co/datasets/facebook/PE-Video)  
Tier: SSS  
Description: Facebook PE Video contains quality high resolution videos with proper keywords and captioning. The videos feel staged though and contains diverse set of objects.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753591734857/f8f42e12-8736-4305-94c8-a8448c781cf6.png align="center")

---

### LLaVA-Video-178K

Link: [lmms-lab/LLaVA-Video-178K](https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K)  
Tier: SS  
Description: Contains various videos with a nice distribution on various types of videos.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753592509289/eda38087-eba8-45b6-a008-bbfbe3a24e24.png align="center")

---

### ByteDance Synthetic Videos

Link: [kevinzzz8866/ByteDance\_Synthetic\_Videos](https://huggingface.co/datasets/kevinzzz8866/ByteDance_Synthetic_Videos)  
Tier: A  
Description: Turn-table style highly detailed 3D renderings of objects. It almost feels like it was designed for creating 3D rendering to later be able to generate 3D models.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753593046921/815e722c-083b-41f0-870b-28def9297eee.png align="center")

---

### Anime Landscape Videos Splited

Link: [svjack/Anime\_Landscape\_Videos\_Splited](https://huggingface.co/datasets/svjack/Anime_Landscape_Videos_Splited)  
Tier: A  
Description: Contains excepts from anime videos mostly high quality background scenes with delicate and visually pleasing movements.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753593778799/67bb9911-3efe-4f59-a1df-eb4b1fc10094.png align="center")

---

### anime-giph

Link: [nev/anime-giph](https://huggingface.co/datasets/nev/anime-giph)  
Tier: A  
Description: Short excepts from popular animated films. Contains interesting motions and special effects. High quality.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753594722341/f036f882-94c5-46e9-95fc-a4c51d8c7834.png align="center")

---

### Anime Segmentation

Link: [skytnt/anime-segmentation](https://huggingface.co/datasets/skytnt/anime-segmentation)  
Tier: A  
Description: Has anime characters segmented out with proper alpha channel. Seems like it is going to be useful for background removals.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753594973324/2c082c97-b098-4f0f-89c2-8e06b5e10a4f.png align="center")

---

### VideoMMMU

Link: [lmms-lab/VideoMMMU](https://huggingface.co/datasets/lmms-lab/VideoMMMU)  
Tier: A  
Description: Peculiar set of video from distinctive domains like art, humanity, science and engineering. Decent number of videos contain powerpoint like transitions.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753592822713/c06c434c-5716-4b49-9ba0-c3907584cbe8.png align="center")

---

### Anime Video and Anime Screenshots

Link: [Fitsd/anime\_video\_and\_anime\_screenshots](https://huggingface.co/datasets/Fitsd/anime_video_and_anime_screenshots)  
Tier: B  
Description: Video recordings of animated films.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753593473306/8ece98bc-aec3-4d6e-a7df-c3ce9d4509ec.png align="center")

---

### Anime Person Detection (Image Data)

Link: [deepghs/anime\_person\_detection](https://huggingface.co/deepghs/anime_person_detection)  
Tier: B  
Description: Contains various anime screenshots and their labeled bounding boxes for humanoids.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753594007608/645dfa7b-ba85-4b0b-8414-cdde65efadd9.png align="center")

---

**VideoFeedback-videos-mp4**

Link: [hexuan21/VideoFeedback-videos-mp4](https://huggingface.co/datasets/hexuan21/VideoFeedback-videos-mp4)  
Tier: A  
Description: Short videos. Relatively well balanced. 37.7k.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753598018832/b5c82f39-d59b-4bd7-9050-0d3e01540f94.png align="center")

---

### WenhaoWang/VidProM

Link: [WenhaoWang/VidProM](https://huggingface.co/datasets/WenhaoWang/VidProM)  
Tier: B  
Description: Contains lots of generated videos from different models.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753598445109/07c5af11-58eb-46af-97f2-0166e7f7436c.png align="center")

---

### zhiqiulin/video\_captioning

Link: [zhiqiulin/video\_captioning](https://huggingface.co/datasets/zhiqiulin/video_captioning)  
Tier: S  
Description: Very good captioned video dataset. Good data balance.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753598649592/0eec0818-2d45-4b81-9448-3548768a83e2.png align="center")

---

### VatsaDev/video-danbooru-final

Link: [VatsaDev/video-danbooru-final](https://huggingface.co/datasets/VatsaDev/video-danbooru-final)  
Tier: B  
Description: Anime videos (100hrs it says)

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753599588167/668a26e9-14a5-4cc8-b863-217ee0fdd968.png align="center")

---

### Nahrawy/VIDIT-Depth-ControlNet-E

Link: [Nahrawy/VIDIT-Depth-ControlNet-E](https://huggingface.co/datasets/Nahrawy/VIDIT-Depth-ControlNet-E)  
Tier: A  
Description: Realistically rendered 3d scenes that look real along with Depth map.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753599261307/1103807a-2129-48cc-8d9c-46af63b7df1e.png align="center")

---

### daiua/video

Link: [daiua/video](https://huggingface.co/datasets/daiua/video)  
Tier: B  
Description: Short-form videos (Chinese mostly) Contains lots of live-action and emotional scenes.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753599045858/00bccd85-798a-4fe3-bced-7280554f740a.png align="center")

---

### VideoHallu

Link: [IntelligenceLab/VideoHallu](https://huggingface.co/datasets/IntelligenceLab/VideoHallu)  
Tier: B  
Description: Generated video from various tools. Good to compare different video gen services.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600000318/5e449eba-de58-4121-a180-4cd236c79574.png align="center")

---

### WenhaoWang/VideoUFO

Link: [WenhaoWang/VideoUFO](https://huggingface.co/datasets/WenhaoWang/VideoUFO)  
Tier: A  
Description: Detailed captions per keyframe. Video dataset. Seems like he has bunch of papers. Would be good look at other datasets from him.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753599792907/7ad1661d-089a-4541-9d33-2038f6191454.png align="center")

---

### youtube\_videos\_11

Link: [2e8konjak/youtube\_videos\_11](https://huggingface.co/datasets/2e8konjak/youtube_videos_11/tree/main)  
Tier: B  
Description: Self explanatory. There are also different compliations like [2e8konjak/youtube\_videos\_12](https://huggingface.co/datasets/2e8konjak/youtube_videos_12) and [2e8konjak/youtube\_videos\_2](https://huggingface.co/datasets/2e8konjak/youtube_videos_2).

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753597228307/3cac0b25-9e96-4a56-809a-135ed9d1cd74.png align="center")

---

### video-reasoning/morse-500-view

Link: [video-reasoning/morse-500-view](https://huggingface.co/datasets/video-reasoning/morse-500-view)  
Tier: A  
Description: Not directly related to spriteDX project, but it can be used to learn visual semantics.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600242326/b125b6b3-46f1-4b24-b240-ac8df9b91ecd.png align="center")

---

### Nagi\_no\_Asukara\_Videos\_Captioned

Link: [svjack/Nagi\_no\_Asukara\_Videos\_Captioned](https://huggingface.co/datasets/svjack/Nagi_no_Asukara_Videos_Captioned)  
Tier: B  
Description: It focuses on one anime but every shot is paired with very detailed caption and scene description.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600428855/3532df47-b5c7-43d3-a8bc-6eab0da59021.png align="center")

### Sprite Animation

Link: [Loacky/sprite-animation](https://huggingface.co/datasets/Loacky/sprite-animation)  
Tier: B  
Description: Retro style sprite animation collected frame by frame. The dataset is not too big.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753594256731/c02cf2ac-23ed-4d03-b6e6-efe9a24fd3ae.png align="center")

---

### Synth-Vid-Detect

Link: [https://huggingface.co/datasets/ductai199x/synth-vid-detect](https://huggingface.co/datasets/ductai199x/synth-vid-detect)  
Tier: A  
Description: Has both synthetic and real video. It can be used to train a model that detects fake from real.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600724464/974a307c-7f52-4c18-aee7-085f4c43e997.png align="center")

  

### video-dataset-disney-organized

Link: [sayakpaul/video-dataset-disney-organized](https://huggingface.co/datasets/sayakpaul/video-dataset-disney-organized)  
Tier: A  
Description: Well detailed captions for different shots of the video.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600552080/e83e5b7d-512b-4320-91e0-621be3aef604.png align="center")

---

### 8k-video-game-dataset

Link: [Kim2091/8k-video-game-dataset](https://huggingface.co/datasets/Kim2091/8k-video-game-dataset)  
Tier: B  
Description: Contains video game footage every 1-5 second interval. Has various games.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753594506177/20f41c98-8d47-4a84-806d-57b75778d486.png align="center")

---

### 2D\_Video\_Game\_Cartoon\_Character\_Sprite-Sheets

Link: [mgane/2D\_Video\_Game\_Cartoon\_Character\_Sprite-Sheets](https://huggingface.co/datasets/mgane/2D_Video_Game_Cartoon_Character_Sprite-Sheets)  
Tier: C  
Description: Contains spritesheets mostly flash style animation for side scroller games. Dataset is small

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753595263602/7ca3e013-efa2-4e6c-9dc0-f4bc0f11ab08.png align="center")

---

### Pexels Videos

Link: [https://huggingface.co/datasets/minh132/pexels-videos](https://huggingface.co/datasets/minh132/pexels-videos)  
Tier: A  
Description: Quality short videos that have natural camera movements.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753596332985/7049b6c8-5315-4302-8e26-60aab9538a43.png align="center")

---

### Tiger Lab VideoFeedback

Link: [TIGER-Lab/VideoFeedback](https://huggingface.co/datasets/TIGER-Lab/VideoFeedback)  
Tier: A  
Description: Diverse set of videos

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753596551034/9465a9fa-5edf-45b2-a62a-7bc0f3e0f4e6.png align="center")

---

### Changli/Ytb\_Video

Link: [Changli/Ytb\_Video](https://huggingface.co/datasets/Changli/Ytb_Video)  
Tier: B  
Description: Youtube videos. Not sure about data balance.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753600908203/9314ee87-6052-48b5-8021-4a9b4120aafb.png align="center")

---

### svjack/Star\_Rail\_Tribbie\_MMD\_Videos\_Splited\_Captioned\_512x384x1\_conn

Link: [svjack](https://huggingface.co/datasets/svjack/Star_Rail_Tribbie_MMD_Videos_Splited_Captioned_512x384x1_conn)  
Tier: C  
Description: MMD video dataset that seems relatively safe. It also contains Alpha mask versions of the frames which would be useful for training **background removal models**.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753601291804/c2886dea-e60f-4114-8886-08e284798150.png align="center")

---

### QimingLi/videos

Link: [QimingLi/videos](https://huggingface.co/datasets/QimingLi/videos)  
Tier: C  
Description: Small video dataset.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753601070498/41926908-52d7-47fc-ba6e-b0c480a2ad83.png align="center")

---

### laion2B-en-aesthetic (Image Data)

Link: [laion/laion2B-en-aesthetic](https://huggingface.co/datasets/laion/laion2B-en-aesthetic)  
Tier: A  
Description: Subset of Laion dataset that are aesthetically pleasing. The caption quality is much better than what we have seen in other places.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753595951177/dd4a2203-ac50-4148-8b64-d3bf10547c2f.png align="center")

---

### novelai-anime-v3-artist-comparison (Image Data)

Link: [novelai-anime-v3-artist-comparison](https://huggingface.co/datasets/deus-ex-machina/novelai-anime-v3-artist-comparison)  
Tier: C  
Description: Mostly generated data based on artist name as a trigger words.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1753592189097/dbfb8254-1f9f-4314-9dd9-2f0f95f1f44c.png align="center")

---

### Steer Clear List

* [https://huggingface.co/datasets/unavailableshorts/Videos](https://huggingface.co/datasets/unavailableshorts/Videos)
    
* [https://huggingface.co/datasets/bitmind/bm-eidon-video](https://huggingface.co/datasets/bitmind/bm-eidon-video)
    

---

**One other thought**: Seedance 1 Pro generates high quality 5s videos relatively fast cheaply. It may be cheaper to generate synthetic dataset instead and use for training.

— Sprited Dev 🌱 Stay safe, Stay cool.
