Fashion ai dataset.

 ModaNet (2018) was groundbreaking but there have been several more recent datasets and models for fashion segmentation and analysis. Here are some notable recent ones:

DeepFashion2 (2023 Update)

  • 491K images with 801K clothing items
  • 13 clothes categories (similar to ModaNet)
  • More detailed annotations including style, occlusion, zoom-in
  • Higher quality annotations and more diverse images
  • Link: https://github.com/switchablenorms/DeepFashion2

VITON-HD (2022)

  • High resolution virtual try-on dataset
  • 13,679 front-view woman/clothing image pairs
  • High quality segmentation masks
  • Particularly good for virtual try-on applications

FashionAI Dataset (2023)

  • From Alibaba
  • Over 180K images
  • Focus on attribute recognition
  • Detailed hierarchical attribute annotations
  • More modern fashion styles and better image quality

LIP (Look Into Person) Dataset (2022 version)

  • 50,000 images with pixel-level annotations
  • 19 semantic human part labels
  • Multiple viewpoints and poses
  • Human parsing focused but includes detailed clothing segmentation

ACGPN Dataset (2023)

  • 40,000 high-resolution person images
  • Detailed parsing maps
  • Semantic segmentation for clothes
  • Focuses on both parsing and virtual try-on

Key Improvements in Recent Datasets:

  1. Higher resolution images
  2. Better annotation quality
  3. More diverse poses and viewpoints
  4. More modern fashion styles
  5. Better handling of occlusion and layering
  6. More detailed attribute annotations
  7. Multi-task annotations (segmentation + attributes + landmarks)

For your specific use case, I would recommend:

  1. DeepFashion2 as your primary dataset - it's the most comprehensive and recent
  2. Augment with ACGPN if you need higher resolution images
  3. Consider FashionAI if you need very detailed attribute recognition

Pedestrian and human attribute dataset.

 

For Pedestrian Detection:

  1. CityPersons - High-quality pedestrian detection dataset with diverse urban scenes from multiple European cities
  2. Caltech Pedestrian Dataset - Contains approximately 250,000 frames with 350,000 bounding boxes and 2,300 unique pedestrians
  3. INRIA Person Dataset - Includes full-body pedestrians in various poses and backgrounds
  4. MOT (Multiple Object Tracking) Dataset - Contains pedestrians in crowded scenes

For Human Attribute Analysis:

  1. RAP (Richly Annotated Pedestrian) Dataset - Over 40 attributes including clothing types, colors, and accessories
  2. PETA Dataset - Large-scale surveillance person attribute dataset with 19,000 images
  3. Market-1501 Attribute Dataset - Contains 27 attributes for clothing and personal items
  4. DeepFashion Dataset - Focuses on clothing items with detailed annotations

Some considerations when choosing a dataset:

  • Make sure to check the license terms for each dataset
  • Consider the image quality and diversity needed for your specific use case
  • Check if the annotations match your requirements (bounding boxes, attributes, etc.)
  • Verify that the dataset size is sufficient for your model training needs

example code in Python using OpenCV and scikit-learn to build a dataset for video classification

refer to code.


..

import cv2
import os
from sklearn.model_selection import train_test_split

# Define the classes for classification
classes = ['class1', 'class2', 'class3']

# Create a list to store the image paths and labels
data = []

# Loop through the videos and extract frames
for cls in classes:
video_dir = f'data/{cls}/'
for filename in os.listdir(video_dir):
video_path = os.path.join(video_dir, filename)
cap = cv2.VideoCapture(video_path)
while cap.isOpened():
ret, frame = cap.read()
if not ret:
break
frame_path = f'frames/{cls}/{filename}_{cap.get(cv2.CAP_PROP_POS_FRAMES)}.jpg'
cv2.imwrite(frame_path, frame)
data.append((frame_path, cls))
cap.release()

# Split the data into training, validation, and test sets
train_val_data, test_data = train_test_split(data, test_size=0.2, stratify=[x[1] for x in data])
train_data, val_data = train_test_split(train_val_data, test_size=0.2, stratify=[x[1] for x in train_val_data])

# Preprocess the images
def preprocess_image(image):
# Resize the image to (224, 224) and normalize the pixel values
image = cv2.resize(image, (224, 224))
image = image.astype('float32') / 255.0
return image

# Load the images and labels into memory
X_train = []
y_train = []
for image_path, label in train_data:
image = cv2.imread(image_path)
image = preprocess_image(image)
X_train.append(image)
y_train.append(label)

X_val = []
y_val = []
for image_path, label in val_data:
image = cv2.imread(image_path)
image = preprocess_image(image)
X_val.append(image)
y_val.append(label)

X_test = []
y_test = []
for image_path, label in test_data:
image = cv2.imread(image_path)
image = preprocess_image(image)
X_test.append(image)
y_test.append(label)

# Train and evaluate the model
# ...

..




In this example, we first define the classes for classification (in this case, 'class1', 'class2', and 'class3'). We then loop through the videos for each class, extract frames from each video, and store the frame paths and labels in a list called data.

Next, we split the data into training, validation, and test sets using train_test_split() from scikit-learn. We stratify the splits to ensure that each set has a proportional number of frames from each class.

We then define a function called preprocess_image() to resize the images to (224, 224) and normalize the pixel values. We load the images and labels into memory for each set using OpenCV, preprocess them using this function, and store them in lists called X_train, y_train, X_val, y_val, X_test, and y_test.

Finally, we can train and evaluate a video classification model using the preprocessed data. The specific code for this step will depend on the model architecture and training strategy you choose to use.


Thank you.