> For the complete documentation index, see [llms.txt](https://medomicslab.gitbook.io/medfl-app-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://medomicslab.gitbook.io/medfl-app-docs/tutorials/simulation/create-pipeline.md).

# Create pipeline

## Introduction

Using the Federated Learning module, you can create and run your own pipelines by simply dragging and dropping nodes onto the open map. The application offers the flexibility to create multiple configurations and run them simultaneously.

In the next video, we will demonstrate how to create your configurations, verify them, and launch the execution.<br>

## Video tutorial

{% embed url="<https://youtu.be/WR5IC0aVMZ8>" %}

### Getting started

To get started with the simulation click on the simulation box, and you will have the simulation page&#x20;

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FxXQZ1L3C1gLbm8CTKirN%2FGroup%2086.png?alt=media&amp;token=5c28862c-8176-425c-bfcc-53e0836975e1" alt=""><figcaption></figcaption></figure>

### Simulation Page Overview

The **Simulation Page** contains several key components that allow users to design, configure, and execute federated learning experiments.

#### 1️⃣ Workspace Section

This section allows you to browse and open all files that have been added or created within your selected workspace.

#### 2️⃣ Available Nodes

This section contains all the nodes that can be used to build a federated learning pipeline.

#### 3️⃣ The Scene

The **Scene** is the open workspace area where you can drag and drop nodes to design your federated learning pipeline visually.

#### 4️⃣ SQLite Configuration

The **SQLite Section** allows you to specify the SQLite database file that will be used during the experiment.

#### 5️⃣ Actions Section

The **Actions Section** provides control over the simulation workflow.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FlhHnIRAWgBYbm9q6ew87%2FGroup%2093.png?alt=media&amp;token=1f5d25a6-3dd9-4891-8a4a-356bbc57d840" alt=""><figcaption></figcaption></figure>

### <i class="fa-1">:1:</i>. SQLite configuration&#x20;

To configure the SQLite database, you simply need to specify a file that will be used as the database for the experiment.

Click on the **“**<mark style="color:$success;">**Select a DB File**</mark>**”** button (4). A pop-up window will appear, allowing you to choose the database file.

You have two options:

* Create a new database file by entering a name for the file.
* Select an existing database file from your workspace, if one is already available.

Once selected, click **Confirm** and wait for the success message, close the pop-up window. The selected file will be used to store experiment results, logs, and metrics.

{% hint style="info" %}
To ensure that everything worked correctly, verify that the warning indicator has disappeared from the button.
{% endhint %}

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FKa5eAMDGYrlBfycXAHsL%2FGroup%2093%20(1).png?alt=media&amp;token=bf44808a-f090-4ccb-bffb-bb8daed1ded6" alt=""><figcaption></figcaption></figure>

### <i class="fa-2">:2:</i>. Dataset node configuration&#x20;

To begin configuring the **Dataset Node**, you first need to place your datasets into the working space of the application, you should see your datasets appearing on the workspace, if not try to refresh the workspcae and it should appear directly.

You can use your operating system’s file manager to copy the files directly there is no need to upload them through the application.

After placing the datasets in the correct folder:

1. Drag and drop the **Dataset Node** into the **Initialization Block**.
2. Click on the node to configure it.
3. Select your dataset.
4. Make sure to choose the target variable.

This ensures that the dataset is correctly configured for the federated learning experiment.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FqnO51FH45leeRmkx43MF%2FGroup%2093%20(3).png?alt=media&amp;token=2cab445e-2ad2-4c79-ac36-e2eedc91024f" alt=""><figcaption></figcaption></figure>

### <i class="fa-3">:3:</i>. Network configuration&#x20;

The network in the simulation is created manually. We create and configure each client individually. To do this, start by adding a **Network node** and linking it to the **Dataset node**.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2F7t2QRnQg8hVY4VyQ7jxa%2FGroup%2093%20(4).png?alt=media&amp;token=68ff5484-a655-498a-b58a-9b3388a971e7" alt=""><figcaption></figcaption></figure>

Once the node is added, click on it to open a new window where you can create your network.

To create a network, drag and drop one **Server** and the desired number of **Clients**, then connect each client to the central server.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FnCtccaRvgwdW7HqYpSwq%2FGroup%2093%20(5).png?alt=media&amp;token=99c826d6-9caf-4156-a24e-32c63435e010" alt=""><figcaption></figcaption></figure>

After adding the clients and connecting them to the server, configure the server by specifying:

* The number of federated learning rounds
* Whether to activate **Differential Privacy**

For the clients, you must assign a datset to each one and specify whether it is a **Train** or **Test** node.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FtQXQf3rQqPiiI2S29yyq%2FGroup%2094.png?alt=media&amp;token=5d27d52b-ff66-4baa-8ae5-0f5ba9cefa6a" alt=""><figcaption></figcaption></figure>

### <i class="fa-4">:4:</i>. Model configuration&#x20;

To create and add a model to the pipeline, add a **Model node**. This node allows you to use one of two model types: **neural network models** or **XGBoost models**.

#### **1. Neural Network Models**

For neural network models, MEDfl allows you to enable or disable transfer learning.

**Transfer Learning Enabled**

When transfer learning is enabled, the Model node requires the following parameters:

* A pretrained model saved as a `.pkl` file
* The optimizer
* The learning rate
* The prediction threshold
* The number of local epochs : represents the number of local training iterations performed by each client during every federated learning round.

**Transfer Learning Disabled**

When transfer learning is disabled, the Model node allows you to create a new model by specifying the following parameters:

* The model type
* The number of layers
* The hidden layer size
* The optimizer
* The learning rate

{% hint style="info" %}
Currently, MEDfl supports only binary classification models. Support for additional model types is planned for future releases.
{% endhint %}

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FVGSvLMODaQYxTALV23NY%2FGroup%20120.png?alt=media&amp;token=c9819d5a-1a02-46ea-a548-d5b9c972981f" alt=""><figcaption></figcaption></figure>

####

2. #### **XGBoost Models**

MEDfl integrates XGBoost models and allows them to be trained within a federated learning network.

To use an XGBoost model, change the model type to **XGBoost**. The Model node will then display a new set of XGBoost-specific parameters.

* **`xgb_mode`**: Specifies the federated XGBoost training mode. Currently, MEDfl supports the **bagging** mode. Support for the **cyclic** mode is planned for a future release.

{% hint style="info" %}
For more information about the bagging and cyclic aggregation modes, refer to the corresponding documentation.
{% endhint %}

* **`xgb_eval_metric`**: Specifies the evaluation metric used to assess the model's performance.
* **`xgb_local_num_boost_rounds`**: Specifies the number of trees trained and added by each client during every federated learning round.
* **`xgb_max_depth`**: Specifies the maximum depth of each decision tree.
* **`xgb_eta`**: Specifies the learning rate used during model training.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FLosatWx6lXGAWznXdwg6%2FGroup%20120%20(2).png?alt=media&amp;token=8d869763-b8ba-4571-b9f1-2ffe4d240a24" alt=""><figcaption></figcaption></figure>

### <i class="fa-5">:5:</i>. Federated learning strategy&#x20;

The pipeline system of the MEDfl application allows you to create multiple pipelines in one session. In our experiment, we will test **two pipelines**.

To do this, we add two **Strategy nodes**, where each strategy uses a different aggregation function.

Start by adding one Strategy node and configure it as follows:

* **Aggregation strategy** FedAvg
* **Evaluation fraction = 1** → 100% of clients are used for evaluation
* **Training fraction = 1** → 100% of clients are used for training
* **Minimum number of clients = 3** → The server waits for all 3 clients before performing aggregation

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2F5B2dSqFyCFvCuBXgmJff%2FGroup%2094%20(3).png?alt=media&amp;token=481ea940-5bb0-4c51-937c-5ee777d946a6" alt=""><figcaption></figcaption></figure>

To create the second Strategy node, use the **Copy Node** feature to duplicate the existing node. Then, modify only the aggregation function.

### <i class="fa-6">:6:</i>. Federated SHAP&#x20;

At the end of the pipeline configuration, you can add a **Federated SHAP node** to calculate the contribution of each feature to the model predictions.

Federated SHAP allows MEDfl to compute feature importance in a federated setting without requiring clients to share their raw data. The local SHAP values calculated by the participating clients are used to produce a federated explanation of the trained model.

To configure the Federated SHAP node, the following parameters are available:

* **`include_client_results`**: Specifies whether the local SHAP results from each client should also be returned.

  When this option is enabled, MEDfl returns both:

  * The **federated SHAP results**, representing the aggregated feature contributions across the federated network.
  * The **local SHAP results** calculated independently by each participating client.

  When this option is disabled, only the federated SHAP results are returned.
* **`advanced_shap_configuration`**: Enables additional configuration options for controlling how SHAP values are calculated.

  When this option is enabled, five additional parameters become available:

  * **`explainer`**: Specifies the SHAP explainer used to calculate feature contributions. The available explainers depend on the model type:
    * For **neural network models**, MEDfl supports **GradientExplainer** and **DeepExplainer**.
    * For **XGBoost models**, MEDfl supports **TreeExplainer**.
  * **`data_split`**: Specifies which local dataset split should be used for the SHAP computation. The available options are:
    * **Training set**
    * **Validation set**
    * **Test set**
  * **`background_size`**: Specifies the number of samples used as the background or reference dataset by the SHAP explainer. The background dataset represents the baseline against which feature contributions are calculated.
  * **`explanation_size`**: Specifies the number of samples for which SHAP explanations are generated. Increasing this value provides explanations for more samples but also increases computation time.
  * **`maximum_samples`**: Defines the maximum number of samples that can be used during the SHAP computation. This parameter helps limit the computational and memory cost when working with large datasets.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FvNUbEGh3bcqILRVa4Nsp%2FGroup%20121.png?alt=media&amp;token=ae60e319-2a12-46f2-800e-1035e4c59fb3" alt=""><figcaption></figcaption></figure>

### <i class="fa-7">:7:</i>. Checking and running pipelines&#x20;

After configuring your pipelines, click on the **Run** button.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2F1KDsysQkKjEyriXBuMsi%2FGroup%2094%20(4).png?alt=media&amp;token=313f9fb3-c2d0-43b2-9149-e47f303b3484" alt=""><figcaption></figcaption></figure>

A popup window will appear showing:

* The configurations
* The number of pipelines

You can check the configuration of each pipeline by switching between the tabs.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2F0dnQDUelMKHoSzyDGhNN%2FGroup%2094%20(5).png?alt=media&amp;token=eb8ffeb7-c9cd-4c9d-8ee7-bd50264fdc55" alt=""><figcaption></figcaption></figure>

Click on **Run Pipeline**, and monitor the execution progress using the progress bar.

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FeyoU3e1vxR68JGrQxMOS%2FGroup%2094%20(6).png?alt=media&amp;token=f88d664e-11f5-40ca-904c-81ac30f69ac3" alt=""><figcaption></figcaption></figure>

### <i class="fa-8">:8:</i>. Checking and running pipelines&#x20;

After the training is completed, click on **See Results** to display the results.

In the results section, you can:

1. Switch between different pipeline configurations
2. View global results
3. View results per client
4. Compare results across configurations

For more details about all the available inter-retations and results see the [Pipeline results](/medfl-app-docs/tutorials/simulation/pipeline-results.md) page&#x20;

<figure><img src="https://2289920470-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FOpYh7tOQkAu0q6hPpqe9%2Fuploads%2FdURrpKd9rtpYSQc1nHMe%2FGroup%20121%20(1).png?alt=media&amp;token=c2a8e479-4cd6-48fb-a7b1-f2a3146add94" alt=""><figcaption></figcaption></figure>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://medomicslab.gitbook.io/medfl-app-docs/tutorials/simulation/create-pipeline.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
