NOTE: This project has been forked from Srijan Das's I3D repository.
The Toyota SmartHome paper (Das et al. 2019). Describes a Separable spatio-temporal attention (STA) neural network consisting of two branches, namely an LSTM and an inflated 3D convolutional neural network (I3D).
The I3D branch takes the action volumes (XYT data, with RGB information), and produces a convolutional feature which is then modulated via an attention block that uses skeletal data (output from an LSTM network).
The way the images for this action volumes is cropped changes what this I3D network sees, and therefore, preprocessing of images can be an interesting alternative to improve performance. The original paper uses an SSD network to crop the images around the detected individual. However, these crops are not provided with the Toyota SmartHome dataset.
In our Sensors paper (Climent-Pérez et al. 2021, https://doi.org/10.3390/s21031005), we have tried to alternatives with success:
- Using a Mask RCNN instead of an SSD.
- Using 'full crops', i.e. full activity crops where information from all detections is used to create a crop of the image containing the full activity (i.e. the resulting activity bounding box contains all detections).
Because the Mask-RCNN, as any other detection network, may fail for some frames during detection, a gap filling
technique has been devised. The figure below shows blue bars for the size of identified gaps, showing that almost half
the gaps are 1 frame long (i.e. most are glitches). Then the vast majority are
below 1 second of video (approx. 20 frames). And that more than 99% are less than 60 frames in length (orange line shows
cumulative count of the gaps, with convergence to 100%).

This information is then used to perform some preprocessing of the raw detections, and fill in gaps below a threshold of 60 frames. To do so, we use the information from the last detection rectangle available.
The preprocessing/ folder contains 6 numbered scripts. Scripts 1 and 2 identify and perform the filling in of the
missing information. These are short descriptions for each of the scripts:
step1_find_gaps.pyis used to determine the existing gaps in detection (i.e. when the Mask RCNN network returned no bounding boxes).step2_fill_gaps.pyis then used to fill in detected gaps, up to a certain size, given certain conditions (i.e. the gaps in detection are not found in the beginning or end of a video).step3_crop_videos.py, can be used to crop the videos using the new 'filled in' information.step4_tar_gz_crops.py, (optional) can be used to create.tgzfiles for the sequences, if necessary.step5_regenerate_splits.py, takes into account sequences for which there are no detections, and changes the splits accordingly.step6_check_crop_images.pycan be used optionally to check whether the generated images are valid in all cases (non-empty). This is useful if the cropping process was interrupted and checks corrupted file contents.
The script in preprocessing/step3_crop_videos.py contains two functions to crop the videos. One is the normal
process of cropping a square are around the detected individual. This is called process_one_video().
Furthermore, there is the option to extract full activity crops by instead calling process_one_video_fullcrop().
Apart from some code clean-up of unused characteristics (non-local, NL), as well as some other scripts. As described in our paper, other changes are aimed at learning rate changes required to train in our case.
- (Das et al. 2019) Das, S., Dai, R., Koperski, M., Minciullo, L., Garattoni, L., Bremond, F., & Francesca, G. (2019). Toyota smarthome: Real-world activities of daily living. In Proceedings of the IEEE International Conference on Computer Vision (pp. 833-842).
- (Climent-Pérez et al. 2021_) Climent-Pérez, P., Florez-Revuelta, F. (2021). Improved action recognition with Separable spatio-temporalattention using alternative Skeletal and Video pre-processing, Sensors 21(3), 1005. DOI: https://doi.org/10.3390/s21031005