Unsupervised Construction of Task-Specific Datasets for Object Re-identification
Autoři
Rok
2021
Publikováno
ICCTA 2021 Conference Proceedings. New York: Association for Computing Machinery, 2021. p. 66-72. ISBN 978-1-4503-9052-1.
Typ
Stať ve sborníku vyzvaná či oceněná
Pracoviště
Anotace
In the last decade, we have seen a significant uprise of deep neural networks in image processing tasks and many other research areas. However, while various neural architectures have successfully solved numerous tasks, they constantly demand more and more processing time and training data. Moreover, the current trend of using existing pre-trained architectures just as backbones and attaching new processing branches on top not only increases this demand but diminishes the explainability of the whole model.
Our research focuses on combinations of explainable building blocks for the image processing tasks, such as object tracking. We propose a combination of Mask R-CNN, state-of-the-art object detection and segmentation neural network, with our previously published method of sparse feature tracking. Such a combination allows us to track objects by connecting detected masks using the proposed sparse feature tracklets. However, this method cannot recover from complete object occlusions and has to be assisted by an object re-identification.
To this end, this paper uses our feature tracking method for a slightly different task: an unsupervised extraction of object representations that we can directly use to fine-tune an object re-identification algorithm. As we have to use objects masks already in the object tracking, our approach utilises the additional information as an alpha channel of the object representations, which further increases the precision of the re-identification. An additional benefit is that our fine-tuning method can be employed even in a fully online scenario.
Video Scene Location Recognition with Neural Networks
Autoři
Rok
2021
Publikováno
Proceedings of the 21st Conference Information Technologies – Applications and Theory (ITAT 2021). Aachen: CEUR Workshop Proceedings, 2021. p. 85-93. vol. 2962. ISSN 1613-0073.
Typ
Stať ve sborníku
Anotace
This paper provides an insight into the possibility of scene recognition from a video sequence with a small set of repeated shooting locations (such as in television series) using artificial neural networks. The basic idea of the presented approach is to select a set of frames from each scene, transform them by a pre-trained single image preprocessing convolutional network, and classify the scene location with subsequent layers of the neural network. The considered networks have been tested and compared on a dataset obtained from The Big Bang Theory television series. We have investigated different neural network layers to combine individual frames, particularly AveragePooling, MaxPooling, Product, Flatten, LSTM, and Bidirectional LSTM layers. We have observed that only some of the approaches are suitable for the task at hand.
Classification Methods for Internet Applications
Autoři
Rok
2020
Publikováno
Cham: Springer, 2020. Studies in Big Data. vol. 69. ISSN 2197-6503. ISBN 978-3-030-36961-3.
Typ
Kniha
Pracoviště
Anotace
This book explores internet applications in which a crucial role is played by classification, such as spam filtering, recommender systems, malware detection, intrusion detection and sentiment analysis. It explains how such classification problems can be solved using various statistical and machine learning methods, including K nearest neighbours, Bayesian classifiers, the logit method, discriminant analysis, several kinds of artificial neural networks, support vector machines, classification trees and other kinds of rule-based methods, as well as random forests and other kinds of classifier ensembles. The book covers a wide range of available classification methods and their variants, not only those that have already been used in the considered kinds of applications, but also those that have the potential to be used in them in the future. The book is a valuable resource for post-graduate students and professionals alike.
Rules extraction from neural networks trained on multimedia data
Autoři
Fanta, M.; Pulc, P.; Holeňa, M.
Rok
2019
Publikováno
Proceedings of the 19th Conference Information Technologies - Applications and Theory (ITAT 2019). Aachen: CEUR Workshop Proceedings, 2019. p. 26-35. vol. 2473. ISSN 1613-0073.
Typ
Stať ve sborníku
Pracoviště
Anotace
Since the universal approximation property ofartificial neural networks was discovered in the late 1980s,i.e., their capability to arbitrarily well approximate nearlyarbitrary relationships and dependences, a full exploitationof this property has been always hindered by the very lowhuman-comprehensibility of the purely numerical repre-sentation that neural networks use for such relationshipsand dependences. The mainstream of attempts to miti-gate that incomprehensibility are methods extracting, fromthe numerical representation, rules of some formal logic,which are in general viewed as human-comprehensible.Many dozens of such methods have already been proposedsince the 1980s, differing in a number of diverse aspects.Due to that diversity, and also due to a close connectionof the semantics of extracted rules to the repsective ap-plication domain, no rules extraction methods have everbecome a standard, and it is always necessary to select asuitable method for the considered domain. Here, rulesextraction from trained neural networks is employed formultimedia data, which is an increasingly important butalso increasingly complex kind of data. Three particularrules extraction methods are considered and applied to themodalities recognized text data and the speech acousticdata, both of them with different subsets of features. Adetailed comparison of the performance of the consideredmethods on those datasets is presented, and a statisticalanalysis of the obtained results is performed.
Motion Segmentation by Semi-Supervised Classification in Dynamic Scenery
Autoři
Pulc, P.; Keruľ-Kmec, O.; Šabata, T.; Holeňa, M.
Rok
2018
Publikováno
Proceedings of Poster Session of 3rd ECML/PKDD Workshop on Advanced Analytics and Learning on Temporal Data (AALTD 2018). Diblin: University College Dublin, 2018. p. 65-72.
Typ
Stať ve sborníku
Pracoviště
Anotace
Automatic description of multimedia content heavily relies
on the ability to discover a structure in such data. As our current focus
is given to efficient multimedia indexing, we are mainly interested in
discovery and segmentation of foreground objects from the background
and their respective description or classification.
Although many approaches based on Convolution Neural Networks have
emerged lately, they are usually executed on all frames from the me-
dia separately which is, to our belief, wasteful and poorly scalable. On
the other hand, methods based on visual Simultaneous Localisation and
Mapping (visual SLAM) utilise the temporal structure of the motion
picture to extract at first a model of the environment and objects in the
scene and later pass these models to methods for object description.
In this paper, we will discuss the first two parts of the visual SLAM –
motion tracking and segmentation. While many approaches impose strict
restrictions in the segmentation phase to filter motion tracking outliers,
we introduce restrictions to the motion tracking itself. Such approach
enables us to use of-the-shelf semi-supervised classification methods in
the motion segmentation phase without explicit outlier filtering.
Semi-supervised and Active Learning in Video Scene Classification from Statistical Features
Autoři
Šabata, T.; Pulc, P.; Holeňa, M.
Rok
2018
Publikováno
Proceedings of the Workshop on Interactive Adaptive Learning (IAL 2018) co-located with European Conference on Machine Learning (ECML 2018) and Principles and Practice of Knowledge Discovery in Databases (PKDD 2018). Aachen: CEUR Workshop Proceedings, 2018. p. 24-35. ISSN 1613-0073.
Typ
Stať ve sborníku
Pracoviště
Anotace
In multimedia classification, the background is usually con-
sidered an unwanted part of input data and is often modeled only to be
removed in later processing. Contrary to that, we believe that a back-
ground model (i.e., the scene in which the picture or video shot is taken)
should be included as an essential feature for both indexing and follow-
up content processing. Information about image background, however,
is not usually the main target in the labeling process and the number of
annotated samples is very limited.
Therefore, we propose to use a combination of semi-supervised and active
learning to improve the performance of our scene classifier, specifically
a combination of self-training with uncertainty sampling. As a result,
we utilize a combination of statistical features extractor, a feed-forward
neural network and support vector machine classifier, which consistently
achieves higher accuracy on less diverse data. With the proposed ap-
proach, we are currently able to achieve precision over 80% on a dataset
trained on a single series of a popular TV show.
Semisupervised segmentation of UHD video
Autoři
Keruľ-Kmec, O.; Pulc, P.; Holeňa, M.
Rok
2018
Publikováno
Proceedings of the 18th Conference Information Technologies - Applications and Theory (ITAT 2018). Aachen: CEUR Workshop Proceedings, 2018. p. 100-107. vol. 2203. ISSN 1613-0073. ISBN 9781727267198.
Typ
Stať ve sborníku
Pracoviště
Anotace
One of the key preprocessing tasks in informa-
tion retrieveal from video is the segmentation of the scene,
primarily its segmentation into foreground objects and the
background. This is actually a classification task, but with
the specific property that it is very time consuming and
costly to obtain human-labelled training data for classifier
training. That suggests to use semisupervised classifiers to
this end. The presented work in progress reports the inves-
tigation of semisupervised classification methods based on
cluster regularization and on fuzzy c-means in connection
with the foreground / background segmentation task. To
classify as many video frames as possible using only a single human-based frame, the semisupervised classifica-
tion is combined with a frequently used keypoint detec-
tor based on a combination of a corner detection method
with a visual descriptor method. The paper experimentally
compares both methods, and for the first of them, also clas-
sifiers with different delays between the human-labelled
video frame and classifier training.
Sentiment analysis from utterances
Autoři
Kožusznik, J.; Pulc, P.; Holeňa, M.
Rok
2018
Publikováno
Proceedings of the 18th Conference Information Technologies - Applications and Theory (ITAT 2018). Aachen: CEUR Workshop Proceedings, 2018. p. 92-99. vol. 2203. ISSN 1613-0073. ISBN 9781727267198.
Typ
Stať ve sborníku
Anotace
The recognition of emotional states in speech is
starting to play an increasingly important role. However,
it is a complicated process, which heavily relies on the
extraction and selection of utterance features related to the
emotional state of the speaker. In the reported research,
MPEG-7 low level audio descriptors[10] serve as features
for the recognition of emotional categories. To this end, a methodology combining MPEG-7 with several important
kinds of classifiers is elaborated.
Towards Real-time Motion Estimation in High-Definition Video Based on Points of Interest
Autoři
Rok
2017
Publikováno
Proceedings of the 2017 Federated Conference on Computer Science and Information Systems. Katowice: Polish Information Processing Society, 2017. p. 67-70. Annals of Computer Science and Information Systems. vol. 11. ISSN 2300-5963. ISBN 978-83-946253-7-5.
Typ
Stať ve sborníku
Pracoviště
Anotace
Currently used motion estimation is usually based on a computation of optical flow from individual images or short sequences. As these methods do not require an extraction of the visual description in points of interest, correspondence can be deduced only by the position of such points.
In this paper, we propose an alternative motion estimation method solely using a binary visual descriptor. By tuning the internal parameters, we achieve either a detection of longer time series or a higher number of shorter series in a shorter time. As our method uses the visual descriptors, their values can be directly used in more complex visual detection tasks.
Application of Meta-learning Principles in Multimedia Indexing
Autoři
Rok
2016
Publikováno
DATESO 2016: Databases, Texts, Specifications, and Objects. Ostrava: Vysoká škola báňská - Technická univerzita Ostrava. Archiv VŠB-TUO, 2016. p. 1-11. ISBN 978-80-248-4031-4.
Typ
Stať ve sborníku
Anotace
Databases of video content traditionally rely on annotations and meta-data imported by a person, usually the uploader. This is supposedly due to a lack of an universal approach to the automated multimedia content annotation. As it may be hard or impossible to find a single classifier for all encountered combinations of different modalities or even a network of the classifiers, current interest of our research is to use meta-learning for multiple stages of the multimedia content classification. With this, we hope to handle correctly all modalities involved including their overlaps. Successively, the extracted classes will be used to build the index and later used for searching and discovery in the multimedia.
Image Processing in Collaborative Open Narrative Systems
Autoři
Pulc, P.; Rosenzveig, E.; Holeňa, M.
Rok
2016
Publikováno
ITAT 2016: Information Technologies - Applications and Theory: Conference on Theory and Practice of Information Technologies. Luxemburg: CreateSpace Independent Publishing Platform, 2016. p. 155-162. ISSN 1613-0073.
Typ
Stať ve sborníku
Pracoviště
Anotace
Open narrative approach enables the creators of multimedia content to create multi-stranded, navigable narrative environments. The viewer is able to navigate such space depending on author’s predetermined constraints, or even browse the open narrative structure arbitrarily based on their interests. This philosophy is used with great advantage in the collaborative open narrative system NARRA. The platform creates a possibility for documentary makers, journalists, activists or other artists to link their own audiovisual material to clips of other authors and finally create a navigable space of individual multimedia pieces.
To help authors focus on building the narratives themselves, a set of automated tools have been proposed. Most obvious ones, as speech-to-text, are already incorporated in the system. However other, more complicated authoring tools, primarily focused on creating metadata for the media objects, are yet to be developed. Most complex of them involve an object description in media (with unrestricted motion, action or other features) and detection of near-duplicates of video content, which is the focus of our current interest.
In our approach, we are trying to use motion-based features and register them across the whole clip. Using GridCut algorithm to segment the image, we then try to select only parts of the motion picture, that are of our interest for further processing. For the selection of suitable description methods, we are developing a meta-learning approach. This will supposedly enable automatic annotation based not only on clip similarity per se, but rather on detected objects present in the shot.
Search for Structure in Audiovisual Recordings of Lectures and Conferences
Autoři
Kopp, M.; Pulc, P.; Holeňa, M.
Rok
2015
Publikováno
ITAT 2015 conference proceedings. Aachen: CEUR Workshop Proceedings, 2015. p. 150-158. ISSN 1613-0073. ISBN 978-1-5151-2065-0.
Typ
Stať ve sborníku
Pracoviště
Anotace
With the quickly rising popularity of multimedia, especially of the audiovisual data, the need to understand the inner structure of such data is increasing. In this case study, we propose a method for structure discovery in recorded lectures. The method consists in integrating a self-organizing map (SOM) and hierarchical clustering to find a suitable cluster structure of the lectures.
The output of every SOM is evaluated by various levels of hierarchical clustering with different number of clusters mapped to the SOM. Within these mapped levels we search for the one with the lowest average within-cluster distance, which we consider the most appropriate number of clusters for the map. In experiments, we applied the proposed approach, with SOMs of four different sizes, to nearly 16 000 slides extracted from the recorded lectures.
Case Study in Approaches to the Classification of Audiovisual Recordings of Lectures and Conferences
Autoři
Rok
2014
Publikováno
Proceedings of the 14th conference ITAT 2014 – Workshops and Posters. Praha: Institute of Computer Science AS CR, 2014. pp. 79-84. ISBN 978-80-87136-19-5.
Typ
Stať ve sborníku
Pracoviště
Anotace
Several methods for classification of semistructured documents already exist, thus also classifications for individual modalities of multimedia content. However, every classifier can behave differently on different data modalities and can be differently appropriate for classification of the considered multimedia content as a whole. Because of that, relying on a single classifier or a static weighting of the classification of individual modalities is not adequate. The present paper describes a case study in searching for suitable classification methods, and in investigating appropriate methods for the aggregation of their results to determine a final class of a lecture or conference recording.