practice_sdoml package

Subpackages

Submodules

practice_sdoml.app module

Deployment and Interactive Demonstrations for SDOML. Built with Gradio, leveraging core modules from practice_sdoml.

practice_sdoml.app.generate_feature_plot(feature_name: str)[source]
practice_sdoml.app.get_data_summary()[source]
practice_sdoml.app.load_dataset_sample()[source]

Loads exploratory dataset sample prioritizing packaged real data.

practice_sdoml.app.run_evaluation()[source]

Evaluates SimpleNet exclusively on the 20% test partition (3,000 samples).

practice_sdoml.app.train_interactive(epochs: int, lr: float, batch_size: int, progress=<gradio.helpers.Progress object>)[source]

Train SimpleNet uploading Gradio loading bar

practice_sdoml.config module

practice_sdoml.dataset module

Data ingestion and preprocessing module with train/test partitioning support.

class practice_sdoml.dataset.DiabetesDataset(csv_path: Path | str | None = None, split: str = 'train', test_size: float = 0.2, random_state: int = 42)[source]

Bases: Dataset

PyTorch Dataset for diabetes_risk.csv supporting stratified train/test splits.

Parameters:
  • csv_path (Path | str | None) – Optional explicit path to the CSV file.

  • split (str) – Target data partition (‘train’, ‘test’, or ‘all’).

  • test_size (float) – Fraction of samples allocated to the test subset.

  • random_state (int) – Random seed for deterministic reproducibility.

practice_sdoml.dataset.get_dataloader(batch_size: int = 32, shuffle: bool = True, split: str = 'train') → Tuple[DataLoader, int, int][source]

Factory helper function to instantiate DiabetesDataset and build a DataLoader.

Parameters:
  • batch_size (int) – Number of samples per batch.

  • shuffle (bool) – Whether to shuffle samples at every epoch.

  • split (str) – Partition identifier (‘train’, ‘test’, or ‘all’).

Returns:

Tuple containing (DataLoader, num_features, num_classes).

Return type:

Tuple[DataLoader, int, int]

practice_sdoml.features module

practice_sdoml.features.main(input_path: Path = PosixPath('/home/runner/work/practice_sdoml_jn/practice_sdoml_jn/data/processed/dataset.csv'), output_path: Path = PosixPath('/home/runner/work/practice_sdoml_jn/practice_sdoml_jn/data/processed/features.csv'))[source]

practice_sdoml.plots module

Graphic utilities for model evaluation module

practice_sdoml.plots.plot_calibration_curve(targets: ndarray, probs: ndarray, num_classes: int, output_path: Path) → None[source]

Creates and saves fiability diagram.

Parameters:
  • targets – True labels.

  • probs – Predicted probabilities.

  • num_classes – Class number.

  • output_path – FIle route for saving images.

practice_sdoml.plots.plot_confusion_matrix(targets: ndarray, preds: ndarray, output_path: Path) → None[source]

Creates and saves confusion matrix

Parameters:
  • targets – True labels.

  • preds – Model predictions.

  • output_path – FIle route for saving images.

practice_sdoml.plots.plot_top_loss_samples(sample_losses: ndarray, output_path: Path, top_k: int = 10) → None[source]

Creates a barplot with the worst classifications.

Parameters:
  • sample_losses – Loss array per sample.

  • output_path – FIle route for saving images.

  • top_k – Sample number to show.

Module contents

Paquete principal de practice_sdoml.

class practice_sdoml.DiabetesDataset(csv_path: Path | str | None = None, split: str = 'train', test_size: float = 0.2, random_state: int = 42)[source]

Bases: Dataset

PyTorch Dataset for diabetes_risk.csv supporting stratified train/test splits.

Parameters:
  • csv_path (Path | str | None) – Optional explicit path to the CSV file.

  • split (str) – Target data partition (‘train’, ‘test’, or ‘all’).

  • test_size (float) – Fraction of samples allocated to the test subset.

  • random_state (int) – Random seed for deterministic reproducibility.

class practice_sdoml.SimpleNet(input_dim: int, num_classes: int)[source]

Bases: Module

Feedfordward neural network for tabular classification.

This architecture is created of intermediate lineal layers combined with non-linear activation functions (ReLU) to process input characteristics and predict class probability.

Parameters:
  • input_dim (int) – Feature number.

  • num_classes (int) – Total number of classes to predict.

forward(x)[source]

Executes fordwars pass of the neural network.

It takes an input tensor batch and spread it over the sequential layers in order to generate non-normalized logits.

Parameters:

x (torch.Tensor) – Input tensor with dimensions (batch_size, input_dim)

Returns:

Output logits with dimensions (batch_size, num_classes)

Return type:

torch.Sensor

practice_sdoml.get_dataloader(batch_size: int = 32, shuffle: bool = True, split: str = 'train') → Tuple[DataLoader, int, int][source]

Factory helper function to instantiate DiabetesDataset and build a DataLoader.

Parameters:
  • batch_size (int) – Number of samples per batch.

  • shuffle (bool) – Whether to shuffle samples at every epoch.

  • split (str) – Partition identifier (‘train’, ‘test’, or ‘all’).

Returns:

Tuple containing (DataLoader, num_features, num_classes).

Return type:

Tuple[DataLoader, int, int]

practice_sdoml.train(epochs: int = 3, lr: float = 0.001)[source]

Trains SimpleNet model using provided data by DiabetesDataset