Data Package (ria_toolkit_oss.data)
The Data package contains abstract data types tailored for radio machine learning, such as Recording, as well
as the abstract interfaces for the radio dataset and radio dataset builder framework.
- class ria_toolkit_oss.data.Annotation(sample_start, sample_count, freq_lower_edge, freq_upper_edge, label='', comment='', detail=None)[source]
Bases:
objectSignal annotations are labels or additional information associated with specific data points or segments within a signal. These annotations could be used for tasks like supervised learning, where the goal is to train a model to recognize patterns or characteristics in the signal associated with these annotations.
Annotations can be used to label interesting points in your recording.
- Parameters:
sample_start (int) – The index of the starting sample of the annotation.
sample_count (int) – The index of the ending sample of the annotation, inclusive.
freq_lower_edge (float) – The lower frequency of the annotation.
freq_upper_edge (float) – The upper frequency of the annotation.
label (str, optional) – The label that will be displayed with the bounding box in compatible viewers including IQEngine. Defaults to an emtpy string.
comment (str, optional) – A human-readable comment. Defaults to an empty string.
detail (dict, optional) – A dictionary of user defined annotation-specific metadata. Defaults to None.
- is_valid()[source]
Verify
sample_count > 0and thefreq_lower_edge < freq_upper_edge.- Returns:
True if valid, False if not.
- overlap(other)[source]
Quantify how much the bounding box in this annotation overlaps with another annotation.
- Parameters:
other (Annotation) – The other annotation.
- Returns:
The area of the overlap in samples*frequency, or 0 if they do not overlap.
- class ria_toolkit_oss.data.Recording(data, metadata=None, dtype=None, timestamp=None, annotations=None)[source]
Bases:
objectTape of complex IQ (in-phase and quadrature) samples with associated metadata and annotations.
Recording data is a complex array of shape C x N, where C is the number of channels and N is the number of samples in each channel.
Metadata is stored in a dictionary of key value pairs, to include information such as sample_rate and center_frequency.
Annotations are a list of
Annotation, defining bounding boxes in time and frequency with labels and metadata.Here, signal data is represented as a NumPy array. This class is then extended in the RIA Backends to provide support for different data structures, such as Tensors.
Recordings are long-form tapes can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, machine learning datasets are curated from collection of recordings by segmenting these longer-form tapes into shorter units called slices.
All recordings are assigned a unique 64-character recording ID,
rec_id. If this field is missing from the provided metadata, a new ID will be generated upon object instantiation.- Parameters:
data (array_like) – Signal data as a tape IQ samples, either C x N complex, where C is the number of channels and N is number of samples in the signal. If data is a one-dimensional array of complex samples with length N, it will be reshaped to a two-dimensional array with dimensions 1 x N.
metadata (dict, optional) – Additional information associated with the recording.
annotations (list of Annotations, optional) – A collection of
Annotationobjects defining bounding boxes.dtype (float or int, optional) – Explicitly specify the data-type of the complex samples. Must be a complex NumPy type, such as
numpy.complex64ornumpy.complex128. Default is None, in which case the type is determined implicitly. Ifdatais a NumPy array, the Recording will use the dtype ofdatadirectly without any conversion.timestamp – The timestamp when the recording data was generated. If provided, it should be a float or integer representing the time in seconds since epoch (e.g.,
time.time()). Only used if the timestamp field is not present in the provided metadata.
- Raises:
ValueError – If data is not complex 1xN or CxN.
ValueError – If metadata is not a python dict.
ValueError – If metadata is not json serializable.
ValueError – If annotations is not a list of valid annotation objects.
Examples:
>>> import numpy >>> from ria_toolkit_oss.data import Recording, Annotation
>>> # Create an array of complex samples, just 1s in this case. >>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> # Create a dictionary of relevant metadata. >>> sample_rate = 1e6 >>> center_frequency = 2.44e9 >>> metadata = { ... "sample_rate": sample_rate, ... "center_frequency": center_frequency, ... "author": "me", ... }
>>> # Create an annotation for the annotations list. >>> annotations = [ ... Annotation( ... sample_start=0, ... sample_count=1000, ... freq_lower_edge=center_frequency - (sample_rate / 2), ... freq_upper_edge=center_frequency + (sample_rate / 2), ... label="example", ... ) ... ]
>>> # Store samples, metadata, and annotations together in a convenient object. >>> recording = Recording(data=samples, metadata=metadata, annotations=annotations) >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': 'me'} >>> print(recording.annotations[0].label) 'example'
- property data
- Returns:
Recording data, as a complex array.
- Type:
Note
For recordings with more than 1,024 samples, this property returns a read-only view of the data.
Note
To access specific samples, consider indexing the object directly with
rec[c, n].
- property dtype
- Returns:
Data-type of the data array’s elements.
- Type:
numpy dtype object
- property sample_rate
- Returns:
Sample rate of the recording, or None if ‘sample_rate’ is not in metadata.
- Type:
- astype(dtype)[source]
Copy of the recording, data cast to a specified type.
- Parameters:
dtype (NumPy data type, optional) – Data-type to which the array is cast. Must be a complex scalar type, such as
numpy.complex64ornumpy.complex128.
- Returns:
A new recording with the same metadata and data, with dtype.
Examples:
Todo
Usage examples coming soon!
- add_to_metadata(key, value)[source]
Add a new key-value pair to the recording metadata.
- Parameters:
key (str) – New metadata key, must be snake_case.
value (any) – Corresponding metadata value.
- Raises:
ValueError – If key is already in metadata or if key is not a valid metadata key.
ValueError – If value is not JSON serializable.
- Returns:
None.
Examples:
Create a recording and add metadata:
>>> import numpy >>> from ria_toolkit_oss.data import Recording >>> >>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { >>> "sample_rate": 1e6, >>> "center_frequency": 2.44e9, >>> } >>> >>> recording = Recording(data=samples, metadata=metadata) >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'timestamp': 17369..., 'rec_id': 'fda0f41...'} >>> >>> recording.add_to_metadata(key="author", value="me") >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': 'me', 'timestamp': 17369..., 'rec_id': 'fda0f41...'}
- update_metadata(key, value)[source]
Update the value of an existing metadata key, or add the key value pair if it does not already exist.
- Parameters:
key (str) – Existing metadata key.
value (any) – New value to enter at key.
- Raises:
ValueError – If value is not JSON serializable
ValueError – If key is protected.
- Returns:
None.
Examples:
Create a recording and update metadata:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { >>> "sample_rate": 1e6, >>> "center_frequency": 2.44e9, >>> "author": "me" >>> }
>>> recording = Recording(data=samples, metadata=metadata) >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': "me", 'timestamp': 17369... 'rec_id': 'fda0f41...'}
>>> recording.update_metadata(key="author", value=you") >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': "you", 'timestamp': 17369... 'rec_id': 'fda0f41...'}
- remove_from_metadata(key)[source]
Remove a key from the recording metadata. Does not remove key if it is protected.
- Parameters:
key (str) – The key to remove.
- Raises:
ValueError – If key is protected.
- Returns:
None.
Examples:
Create a recording and add metadata:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { ... "sample_rate": 1e6, ... "center_frequency": 2.44e9, ... }
>>> recording = Recording(data=samples, metadata=metadata) >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'timestamp': 17369..., # Example value 'rec_id': 'fda0f41...'} # Example value
>>> recording.add_to_metadata(key="author", value="me") >>> print(recording.metadata) {'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': 'me', 'timestamp': 17369..., # Example value 'rec_id': 'fda0f41...'} # Example value
- view(output_path='images/signal.png', **kwargs)[source]
Create a plot of various signal visualizations as a PNG image.
- Parameters:
output_path (str, optional) – The output image path. Defaults to “images/signal.png”.
kwargs – Keyword arguments passed on to ria_toolkit_oss.view.view_sig.
- Type:
dict of keyword arguments
Examples:
Create a recording and view it as a plot in a .png image:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { >>> "sample_rate": 1e6, >>> "center_frequency": 2.44e9, >>> }
>>> recording = Recording(data=samples, metadata=metadata) >>> recording.view()
- simple_view(**kwargs)[source]
Create a plot of various signal visualizations as a PNG or SVG image.
- Parameters:
kwargs – Keyword arguments passed on to ria_toolkit_oss.view.view_signal_simple.view_simple_sig.
- Type:
dict of keyword arguments
Examples:
Create a recording and view it as a plot in a .png image:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { >>> "sample_rate": 1e6, >>> "center_frequency": 2.44e9, >>> }
>>> recording = Recording(data=samples, metadata=metadata) >>> recording.simple_view()
- to_sigmf(filename=None, path=None, overwrite=False)[source]
Write recording to a set of SigMF files.
The SigMF io format is defined by the SigMF Specification Project
- Parameters:
recording (Recording) – The recording to be written to file.
filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.
path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.
- Raises:
IOError – If there is an issue encountered during the file writing process.
- Returns:
None
- to_npy(filename=None, path=None, overwrite=False)[source]
Write recording to
.npybinary file.- Parameters:
- Raises:
IOError – If there is an issue encountered during the file writing process.
- Returns:
Path where the file was saved.
- Return type:
Examples:
Create a recording and save it to a .npy file:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { >>> "sample_rate": 1e6, >>> "center_frequency": 2.44e9, >>> }
>>> recording = Recording(data=samples, metadata=metadata) >>> recording.to_npy()
- to_wav(filename=None, path=None, target_sample_rate=48000, bits_per_sample=32, overwrite=False)[source]
Write recording to WAV file with embedded YAML metadata.
WAV format uses stereo audio with I (in-phase) in left channel and Q (quadrature) in right channel. Metadata is stored in standard LIST INFO chunks with RF-specific metadata encoded as YAML in the ICMT (comment) field for human readability.
- Parameters:
filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.
path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.
target_sample_rate (int, optional) – Sample rate stored in the WAV header when no sample_rate metadata is present. IQ samples are written without decimation or interpolation. Default is 48000 Hz.
bits_per_sample (int, optional) – Bits per sample (32 for float32, 16 for int16). Default is 32.
overwrite (bool, optional) – Whether to overwrite existing files. Default is False.
- Raises:
IOError – If there is an issue encountered during the file writing process.
- Returns:
Path where the file was saved.
- Return type:
Examples:
Create a recording and save it to a .wav file:
>>> import numpy >>> from ria_toolkit_oss.data import Recording >>> samples = numpy.exp(1j * 2 * numpy.pi * 0.1 * numpy.arange(10000)) >>> metadata = {"sample_rate": 1e6, "center_frequency": 915e6} >>> recording = Recording(data=samples, metadata=metadata) >>> recording.to_wav()
- to_blue(filename=None, path=None, data_format='CI', overwrite=False)[source]
Write recording to MIDAS Blue file format.
MIDAS Blue is a legacy RF file format with a 512-byte binary header. Commonly used with X-Midas and other RF/radar signal processing tools.
- Parameters:
filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.
path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.
data_format (str, optional) – Format code (default ‘CI’ = complex int16). Common formats: ‘CI’ (complex int16), ‘CF’ (complex float32), ‘CD’ (complex float64). Integer formats require the IQ samples to already be scaled within [-1, 1).
overwrite (bool, optional) – Whether to overwrite existing files. Default is False.
- Raises:
IOError – If there is an issue encountered during the file writing process.
- Returns:
Path where the file was saved.
- Return type:
Examples:
Create a recording and save it to a .blue file:
>>> import numpy >>> from ria_toolkit_oss.data import Recording >>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = {"sample_rate": 1e6, "center_frequency": 2.44e9} >>> recording = Recording(data=samples, metadata=metadata) >>> recording.to_blue()
- trim(num_samples, start_sample=0)[source]
Trim Recording samples to a desired length, shifting annotations to maintain alignment.
- Parameters:
- Raises:
IndexError – If start_sample + num_samples is greater than the length of the recording.
IndexError – If sample_start < 0 or num_samples < 0.
- Returns:
The trimmed Recording.
- Return type:
Examples:
Create a recording and trim it:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) >>> metadata = { ... "sample_rate": 1e6, ... "center_frequency": 2.44e9, ... }
>>> recording = Recording(data=samples, metadata=metadata) >>> print(len(recording)) 10000
>>> trimmed_recording = recording.trim(start_sample=1000, num_samples=1000) >>> print(len(trimmed_recording)) 1000
- normalize()[source]
Scale the recording data, relative to its maximum value, so that the magnitude of the maximum sample is 1.
- Returns:
Recording where the maximum sample amplitude is 1.
- Return type:
Examples:
Create a recording with maximum amplitude 0.5 and normalize to a maximum amplitude of 1:
>>> import numpy >>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) * 0.5 >>> metadata = { ... "sample_rate": 1e6, ... "center_frequency": 2.44e9, ... }
>>> recording = Recording(data=samples, metadata=metadata) >>> print(numpy.max(numpy.abs(recording.data))) 0.5
>>> normalized_recording = recording.normalize() >>> print(numpy.max(numpy.abs(normalized_recording.data))) 1
Radio Dataset SubPackage
The Radio Dataset Subpackage defines the abstract interfaces and framework components for the management of machine learning datasets tailored for radio signal processing.
- class ria_toolkit_oss.data.datasets.RadioDataset(source)[source]
Bases:
ABCA radio dataset is an iterable dataset designed for machine learning applications in radio signal processing and analysis. They are a structured collections of examples in a machine learning-ready format, with associated metadata.
This is an abstract interface defining common properties and behavior of radio datasets. Therefore, this class should not be instantiated directly. Instead, it should be subclassed to define specific interfaces for different types of radio datasets. For example, see ria_toolkit_oss.data.datasets.IQDataset, which is a radio dataset subclass tailored for tasks involving the processing of radio signals represented as IQ (In-phase and Quadrature) samples.
- Parameters:
source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.
- property shape
- Returns:
The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.
- Type:
tuple of ints
- property data
Retrieve the data from the source file.
Note
Accessing this property reads all the data from the source file into memory as a NumPy array, which can consume significant amounts of memory and potentially degrade performance. Instead, use the
RadioDatasetclass methods to process and manipulate the dataset source file. You can read individual examples into memory as NumPy arrays by indexing the dataset:RadioDataset[idx].- Returns:
The dataset examples as a single NumPy array.
- Type:
- property metadata
Retrieve the metadata from the source file.
Note
Accessing this property reads all the metadata from the source file into memory as a Pandas DataFrame.
- Returns:
The dataset metadata as a Pandas DataFrame.
- Type:
pd.DataFrame
- property labels
Retrieves the metadata labels from the dataset file.
- Returns:
A list of metadata column headers.
- Return type:
list of strings
Examples:
>>> awgn_builder = AWGN_Builder() >>> awgn_builder.download_and_prepare() >>> ds = awgn_builder.as_dataset(backend="pytorch") >>> print(ds.labels) ['rec_id', 'modulation', 'snr']
- abstract inspect()[source]
Todo
This method is not yet fully conceptualized. Likely, it will wrap some of the functionality in the Dataset Inspector package (dataset_manager.inspector) to produce an image or visualization. However, the Dataset Inspector package is not yet implemented.
- abstract default_augmentations()[source]
Returns a list of default augmentations.
- Returns:
A list of default augmentations.
- Return type:
list of callable
- augment(class_key, augmentations=None, level=1.0, target_size=None, classes_to_augment=None, inplace=False)[source]
Supplement the dataset with new examples by applying various transformations to the pre-existing examples in the dataset.
Todo
This method is currently under construction, and may produce unexpected results.
The process of supplementing a dataset to artificially increase the diversity of examples is called augmentation. Training on augmented data can enhance the generalization and robustness of deep machine learning models. For more information, see A Complete Guide to Data Augmentation.
Metadata for each new example will be identical to the metadata of the pre-existing example from which it was generated. The metadata will be extended to include an ‘augmentation’ column, populated with the string representation of the transform used.
Augmented data should only be used for model training, not for testing or validation.
Unless specified, augmentations are applied equally across classes, maintaining the original class distribution.
If target_size does not match the sum of the original class sizes scaled by an integer multiple, the class distribution is slightly adjusted to satisfy target_size.
- Parameters:
class_key (str) – Class name used to augment from and calculate class distribution.
augmentations (callable or list of callables, optional) – A function or list of functions that take an example and return a transformed version. Defaults to
default_augmentations().level (float or list of floats, optional) –
The extent of augmentation from 0.0 (none) to 1.0 (full). If
classes_to_augmentis specified, can be either:A single float: All classes augmented evenly to this level.
A list of floats: Each element corresponds to the augmentation level target for the corresponding class.
target_size (int or list of ints, optional) –
Target size of the augmented dataset. Overrides
levelif specified. Ifclasses_to_augmentis specified, can be either:A single float: All classes are augmented proportional to their relative frequency until the dataset reaches target_size.
A list of floats: Each element corresponds to the target size for the corresponding class.
classes_to_augment (string or list of strings, optional) – List of metadata keys of classes to augment.
inplace (bool, optional) – If True, the augmentation is performed inplace and
Noneis returned.
- Raises:
ValueError – If level has any values not in the range (0,1].
ValueError – If target_size of dataset is already sufficed.
ValueError – If a class in classes_to_augment does not exist in class_key.
- Returns:
The augmented dataset or None if
inplace=True.- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes(class_key='col') {'a': 100, 'b': 500, 'c': 300} >>> new_ds = ds.augment(class_key='col', classes_to_augment=['a', 'b'], target_size=1200) >>> new_ds.get_class_sizes(class_key='col') {'a': 150, 'b': 750, 'c': 300}
- subsample(class_key, percentage, inplace=False)[source]
Reduces the number of examples in all classes of a dataset by randomly subsampling each class according to a specified percentage. This function reduces the number of examples per class to the specified percentage without affecting the overall class distribution.
- Parameters:
class_key (str) – The name of the class to subsample.
percentage (float) – The percentage of the original class sizes to keep.
inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset object and corresponding source file, leaving the original dataset unchanged. Default is False.
- Raises:
ValueError – If the target size of the class with the lowest frequency goes to 0.
- Returns:
The subsampled dataset.
- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes(class_key="col") {a:100, b:200, c:300} >>> new_ds = ds.subsample(percentage=0.80, class_key="col") >>> new_ds.get_class_sizes(class_key="col") {a:80, b:160, c:240}
- resample(quantity_target, class_key, inplace=False)[source]
Adjusts an unsampled dataset by changing the number of examples per class to a user-specified quantity.
- For each class:
If there are excess examples, it randomly subsamples the class to the quantity target.
If there are less examples, it randomly duplicates examples to reach the quantity target.
- Parameters:
quantity_target (int) – The number of examples each class should have.
class_key (str) – The label of the class to resample.
inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset object and corresponding source file, leaving the original dataset unchanged. Default is False.
- Returns:
The resampled dataset.
- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes(class_key="col") {a:100, b:200, c:300} >>> new_ds = ds.resample(quantity_target=250, class_key="col") >>> new_ds.get_class_sizes(class_key="col") {a:250, b:250, c:250}
- homogenize(class_key, example_limit=None, inplace=False)[source]
- Discards excess samples by randomly subsampling all classes within a dataset that have more than a
user-specified limit of examples. If the user doesn’t specify a limit, the class the with the fewest examples is selected as the limit.
- Parameters:
class_key (str) – The label of the class to homogenize.
example_limit (int, optional) – The class size limit to which all classes are subsampled. If not specified, the class with the fewest examples is used as the limit. Default is None.
inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.
- Returns:
The homogenized dataset.
- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes(class_key="col") {a:1000, b:5000, c:1500, d:900} >>> new_ds = ds.homogenize(example_limit=1000, class_key="col") >>> new_ds.get_class_sizes(class_key="col") {a:1000, b:1000, c:1000, d:900}
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes(class_key="col") {a:1000, b:5000, c:1500, d:900} >>> new_ds = ds.homogenize(class_key="col") >>> new_ds.get_class_sizes(class_key="col") {a:900, b:900, c:900, d:900}
- drop_class(class_key, class_value, inplace=False)[source]
Removes an entire class from the dataset.
- Parameters:
class_key (str) – Class that will have a value dropped from it. Example: ‘signal_type’
class_value (str) – Value of the class to be dropped. Example: ‘LTE’, ‘NR’
inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Defaults to False.
- Raises:
ValueError – If the entered class name does not exist in the dataset.
- Returns:
The dataset without the removed class.
- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.get_class_sizes() {a:100, b:500, c:300} >>> new_ds = ds.drop_class('a') >>> new_ds.get_class_sizes() {b:500, c:300}
- add_label(column_name, data, inplace=False)[source]
Add a new metadata label to the dataset.
Todo
This method is not yet implemented.
- Parameters:
column_name – Name of the new metadata column header.
data – The contents of the new metadata column.
inplace (bool, optional) – If True, the label is added inplace and
Noneis returned. Defaults to False.
- Raises:
ValueError – If the length of
datais not equal to the length of the dataset.- Returns:
The augmented dataset or None if
inplace=True.- Return type:
RadioDataset or None
Examples:
Todo
Usage examples coming soon.
- get_class_sizes(class_key)[source]
Returns a dictionary containing the sizes of each class in the dataset at the provided key.
- Parameters:
class_key (str) – The class label.
- Raises:
ValueError – If the specified key is not found in the dataset labels.
- Returns:
A dictionary where each key is a distinct class label, and it’s value is the class size.
- Return type:
A dictionary where the keys are strings and the values are integers
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> spectrogram_sensing_builder = AWGN_Builder() >>> spectrogram_sensing_builder.download_and_prepare() >>> ds = spectrogram_sensing_builder.as_dataset(backend="pytorch") >>> ds.get_class_sizes(class_key='signal_type') {'LTE': 900, 'NR': 900, 'LTE_NR': 900}
- delete_example(idx, inplace=False)[source]
Deletes an example and it’s corresponding metadata from the dataset.
- Parameters:
- Returns:
The new dataset or None if
inplace=True.- Return type:
RadioDataset or None
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> spectrogram_sensing_builder = AWGN_Builder() >>> spectrogram_sensing_builder.download_and_prepare() >>> ds = spectrogram_sensing_builder.as_dataset(backend="pytorch") >>> len(ds) 2700 >>> ds = ds.delete_example(idx=34) >>> len(ds) 2699
- append(example, metadata)[source]
Append a single example to the end of the dataset. This operation is performed inplace.
Todo
This method is not yet implemented.
- Parameters:
example (numpy.typing.ArrayLike) – The example to append.
metadata (dict) – The corresponding metadata dictionary.
- Raises:
ValueError – If example does not the same shape and type as rest of the examples in the dataset.
- Returns:
None.
Examples:
Todo
Usage examples coming soon.
- join(ds)[source]
Join or merge together two radio datasets.
Todo
This method is not yet implemented.
Duplicate entries are not removed; they are included.
The examples are not shuffled; examples from
dsare appended at the end.Metadata will be expanded to contain all columns.
- Parameters:
ds (bool) – The dataset to merge together with self. Examples from both datasets must have the same shape.
- Returns:
The combined dataset.
- Return type:
Examples:
Todo
Usage examples coming soon.
- filter(mask, inplace=False)[source]
Filter the dataset using the provided mask.
Todo
This method is not yet implemented.
- Parameters:
mask (array_like) – A boolean mask. Where True, keep the corresponding examples. Where False, discard keep the corresponding examples. The filtering mask is often the result of applying a condition across the elements of the dataset.
inplace (bool, optional) – If True, the filter operation is performed inplace and
Noneis returned. Defaults to False.
- Returns:
The filtered dataset or None if
inplace=True.- Return type:
RadioDataset or None
Examples:
Todo
Usage examples coming soon!
- class ria_toolkit_oss.data.datasets.IQDataset(source)[source]
Bases:
RadioDataset,ABCAn
IQDatasetis aRadioDatasettailored for machine learning tasks that involve processing radiofrequency (RF) signals represented as In-phase (I) and Quadrature (Q) samples.For machine learning tasks that involve processing spectrograms, please use ria_toolkit_oss.data.datasets.SpectDataset instead.
This is an abstract interface defining common properties and behaviour of IQDatasets. Therefore, this class should not be instantiated directly. Instead, it is subclassed to define custom interfaces for specific machine learning backends.
- Parameters:
source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.
- property shape
- IQ datasets are M x C x N, where M is the number of examples, C is the number of channels, N is the length
of the signals.
- Returns:
The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.
- Type:
tuple of ints
- trim_examples(trim_length, keep='start', inplace=False)[source]
Trims all examples in a dataset to a desired length.
- Parameters:
trim_length (int) – The desired length of the trimmed examples.
keep (str, optional) – Specifies the part of the example to keep. Defaults to “start”. The options are: - “start” - “end” - “middle” - “random”
inplace (bool) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.
- Raises:
ValueError – If trim_length is greater than or equal to the length of the examples.
ValueError – If value of keep is not recognized.
ValueError – If specified trim length is invalid for middle index.
- Returns:
The dataset that is composed of shorter examples.
- Return type:
IQDataset
Examples:
>>> from ria.dataset_manager.builders import AWGN_Builder() >>> builder = AWGN_Builder() >>> builder.download_and_prepare() >>> ds = builder.as_dataset() >>> ds.shape (5, 1, 3) >>> new_ds = ds.trim_examples(2) >>> new_ds.shape (5, 1, 2)
- split_examples(split_factor=None, example_length=None, inplace=False)[source]
If the current example length is not evenly divisible by the provided example_length, excess samples are discarded. Excess examples are always at the end of the slice. If the split factor results in non-integer example lengths for the new example chunks, it rounds down.
For example:
Requires either split_factor or example_length to be specified but not both. If both are provided, split factor will be used by default, and a warning will be raised.
- Parameters:
split_factor (int, optional) – the number of new example chunks produced from each original example, defaults to None.
example_length (int, optional) – the example length of the new example chunks, defaults to None.
inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.
- Returns:
A dataset with more examples that are shorter.
- Return type:
Examples:
If the dataset has 100 examples of length 1024 and the split factor is 2, the resulting dataset will have 200 examples of 512. No samples have been discarded.
If the example dataset has 100 examples of length 1024 and the example length is 100, the resulting dataset will have 1000 examples of length 100. The remaining 24 samples from each example have been discarded.
- class ria_toolkit_oss.data.datasets.SpectDataset(source)[source]
Bases:
RadioDataset,ABCA
SpectDatasetis aRadioDatasettailored for machine learning tasks that involve processing radiofrequency (RF) signals represented as spectrograms. This class is integrated with vision frameworks, allowing you to leverage models and techniques from the field of computer vision for analyzing and processing radio signal spectrograms.For machine learning tasks that involve processing on IQ samples, please use ria_toolkit_oss.data.datasets.IQDataset instead.
This is an abstract interface defining common properties and behaviour of IQDatasets. Therefore, this class should not be instantiated directly. Instead, it is subclassed to define custom interfaces for specific machine learning backends.
- Parameters:
source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.
- property shape
Spectrogram datasets are M x C x H x W, where M is the number of examples, C is the number of image channels, H is the height of the spectrogram, and W is the width of the spectrogram.
- Returns:
The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.
- Type:
tuple of ints
- class ria_toolkit_oss.data.datasets.DatasetBuilder[source]
Bases:
ABCAbstract interface for radio dataset builders. These builder produce radio datasets for common and project datasets related to radio science.
This class should not be instantiated directly. Instead, subclass it to define specific builders for different datasets.
- property version
- Returns:
The version identifier of the dataset.
- Type:
Version Identifier
- property latest_version
- Returns:
The version identifier of the latest available version of the dataset, or None if not set.
- Type:
Version Identifier or None
- property license
- Returns:
The dataset license information.
- Type:
- property info
- Returns:
Information about the dataset including the name, author, and version of the dataset.
- Return type:
- abstract download_and_prepare()[source]
Download and prepare the dataset for use as an HDF5 source file.
Once an HDF5 source file has been prepared, the downloaded files are deleted.
- abstract as_dataset(backend)[source]
A factory method to manage the creation of radio datasets.
- Parameters:
backend (str) – Backend framework to use (“pytorch” or “tensorflow”).
Note: Depending on your installation, not all backends may be available.
- Returns:
A new RadioDataset based on the signal representation and specified backend.
- Type:
- ria_toolkit_oss.data.datasets.split(dataset, lengths)[source]
Split a radio dataset into non-overlapping new datasets of given lengths.
Recordings are long-form tapes, which can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, radio datasets are curated from collections of recordings by segmenting these longer-form tapes into shorter units called slices.
For each slice in the dataset, the metadata should include the unique ID of the recording from which the example was cut (‘rec_id’). To avoid leakage, all examples with the same ‘rec_id’ are assigned only to one of the new datasets. This ensures, for example, that slices cut from the same recording do not appear in both the training and test datasets.
This restriction makes it challenging to generate datasets with the exact lengths specified. To get as close as possible, this method uses a greedy algorithm, which assigns the recordings with the most slices first, working down to those with the fewest. This may not always provide a perfect split, but it works well in most practical cases.
This function is deterministic, meaning it will always produce the same split. For a random split, see ria_toolkit_oss.data.datasets.random_split.
- Parameters:
dataset (RadioDataset) – Dataset to be split.
- Param:
lengths: Lengths or fractions of splits to be produced. If given a list of fractions, the list should sum up to 1. The lengths will be computed automatically as
floor(frac * len(dataset))for each fraction provided, and any remainders will be distributed in round-robin fashion.- Returns:
List of radio datasets. The number of returned datasets will correspond to the length of the provided ‘lengths’ list.
- Return type:
list of RadioDataset
Examples:
>>> import random >>> import string >>> import numpy as np >>> import pandas as pd >>> from ria_toolkit_oss.data.datasets import split
First, let’s generate some random data:
>>> shape = (24, 1, 1024) # 24 examples, each of length 1024 >>> real_part, imag_part = numpy.random.randint(0, 12, size=shape), numpy.random.randint(0, 79, size=shape) >>> data = real_part + 1j * imag_part
Then, a list of recording IDs. Let’s pretend this data was cut from 4 separate recordings:
>>> rec_id_options = [''.join(random.choices(string.ascii_lowercase + string.digits, k=256)) for _ in range(4)] >>> rec_id = [numpy.random.choice(rec_id_options) for _ in range(shape[0])]
Using this data and metadata, let’s initialize a dataset:
>>> metadata = pd.DataFrame(data={"rec_id": rec_id}).to_records(index=False) >>> fid = os.path.join(os.getcwd(), "source_file.hdf5") >>> ds = RadioDataset(source=fid)
Finally, let’s do an 80/20 train-test split:
>>> train_ds, test_ds = split(ds, lengths=[0.8, 0.2])
- ria_toolkit_oss.data.datasets.random_split(dataset, lengths, generator=None)[source]
Randomly split a radio dataset into non-overlapping new datasets of given lengths.
Recordings are long-form tapes, which can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, radio datasets are curated from collections of recordings by segmenting these longer-form tapes into shorter units called slices.
For each slice in the dataset, the metadata should include the unique recording ID (‘rec_id’) of the recording from which the example was cut. To avoid leakage, all examples with the same ‘rec_id’ are assigned only to one of the new datasets. This ensures, for example, that slices cut from the same recording do not appear in both the training and test datasets.
This restriction makes it unlikely that a random split will produce datasets with the exact lengths specified. If it is important to ensure the closest possible split, consider using ria_toolkit_oss.data.datasets.split instead.
- Parameters:
dataset (RadioDataset) – Dataset to be split.
generator (NumPy Generator Object, optional.) – Random generator. Defaults to None.
- Param:
lengths: Lengths or fractions of splits to be produced. If given a list of fractions, the list should sum up to 1. The lengths will be computed automatically as
floor(frac * len(dataset))for each fraction provided, and any remainders will be distributed in round-robin fashion.- Returns:
List of radio datasets. The number of returned datasets will correspond to the length of the provided ‘lengths’ list.
- Return type:
list of RadioDataset
- See Also:
ria_toolkit_oss.data.datasets.split: Usage is the same as for
random_split().