Contenuto principale

Data Windowing for Time Series Anomaly Detection

R2026b

Input-data windowing is a fundamental preprocessing technique in time series anomaly detection (TSAD) that captures behavior patterns rather than the behavior of individual points. By separating input data into uniformly sized segments and evaluating each of these segmented patterns as a whole, data windowing is able to capture sequence dependencies and contextual information that are more perceptive than simply evaluating the sequential data point by point.

Anomaly Detection with Time Series Data

Time series data provides unique opportunities for anomaly detection. When examined in isolation, individual data points often lack sufficient context to indicate whether they represent normal or abnormal behavior. But, when evaluated together at a pattern level, the data can reveal more meaningful information, allowing the detection system to learn to recognize these patterns and distinguish between normal variations and genuine anomalies.

For example, an isolated sensor reading might be within a normal range on its own, but when considered alongside its group of neighbors in the time series, the reading might actually signify an anomaly if the value is drifting over time with respect to its neighbors.

Windowing addresses this ambiguity by grouping and evaluating consecutive points together in fixed-size segments to provide the necessary relational context for accurate anomaly assessment.

Applying Windowing Techniques to Machine Learning Algorithms

Most machine learning algorithms require fixed-size input vectors. However, time series data often comes packaged in inputs of varying sizes. Imposing fixed windowing on the input data stream provides a systematic approach for segmenting variable-length sequences into uniform inputs that machine learning models can process. Each window becomes a standardized observation unit, allowing the model to learn the statistical properties and patterns that characterize normal behavior across multiple time scales You can specify these windows to be either non-overlapping or as sliding segments, that is, as successive segments that overlap with the amount that you specify.

The Stride parameter controls the degree of overlap between consecutive windows, which in turn determines the granularity of analysis. This mechanism ensures comprehensive coverage of the time series while allowing for flexible trade-offs between computational efficiency and detection resolution. Overlapping windows provide dense temporal coverage and smooth anomaly scores, while non-overlapping windows offer computational efficiency for real-time applications.

Different types of anomalies require different window lengths for effective detection. Point anomalies can typically be detected in short windows, while contextual anomalies or collective anomalies may require longer windows to capture the full extent of the abnormal pattern. The windowing approach provides the flexibility to adjust the timescale of analysis according to the specific characteristics of the system you are monitoring and the types of anomalies you expect.

Window-Based Feature Extraction

Within each window, the detector algorithm can characterize the data segment through a set of statistical and spectral features that capture its essential properties. The feature extraction process transforms raw time series segments into meaningful representations that highlight relevant characteristics, such as central tendency, variability, trend, and periodicity. These extracted features provide a multidimensional view of the system behavior, enabling more sophisticated anomaly detection than would be possible with raw data alone.

You can also disable feature extraction if you want to use only the raw data. When feature extraction is disabled, the windowing mechanism operates in a direct mode that preserves the raw time series data within each window and restructures the data for model input. In this configuration, the algorithm flattens the multidimensional time series data within each window into a single vector representation.

For instance, consider a window that contains multiple channels of sensor readings across several time steps. If feature extraction is disabled, the windowing software concatenates the data into a continuous vector that linearizes the time and channel dimensions. This approach maintains all the original data values without a statistical summarization, allowing the machine learning model to learn directly from the raw temporal patterns and interchannel relationships present in the windowed segments.