# Pentaho Academy

<h2 align="center">Pentaho Academy</h2>

**The Pentaho Academy serves as the comprehensive training and enablement hub for organizations implementing Pentaho solutions.**

This specialized learning centre provides hands-on workshops, certification programs, and guided learning paths designed to accelerate customer success with the Pentaho platform.

Through a combination of self-paced online modules, instructor-led sessions, and practical labs using the 30-day Pentaho Enterprise Edition, the Academy ensures customers can effectively leverage their data infrastructure investments.

***

### Getting Started <a href="#getting-started" id="getting-started"></a>

To begin your journey with us, you will need to download and install Pentaho Enterprise - 30 day trial:

{% stepper %}
{% step %}

#### Download

Get the 30‑day Pentaho Enterprise trial:

👀Your License Key is already activated with the 30-day Pentaho Enterprise Edition . This License Key is applicable for any of the supported OS.

<p align="center"><a href="https://pentaho.com/download/" class="button primary" data-icon="arrow-down-from-bracket">Download 30-day Pentaho Enterprise</a></p>
{% endstep %}

{% step %}

#### Install

Follow the Quick Start guide for a streamlined setup on Windows 11:&#x20;

<p align="center"> <a href="https://academy.pentaho.com/pentaho-data-integration/setup/pentaho-lab#pentaho-enterprise-edition" class="button primary" data-icon="book-open">Pentaho Enterprise - Quick Start</a></p>
{% endstep %}

{% step %}

#### Verify your install

Launch Pentaho Server & Data Integration (GUI):

{% tabs %}
{% tab title="Pentaho Server" %}
Windows:

```powershell
cd \
cd Pentaho/server/pentaho-server
./start-server.bat
```

Linux:

```powershell
cd \
cd /opt/pentaho/pentaho-server
./start-server.sh
```

{% endtab %}

{% tab title="Data Integration" %}
Windows:

```powershell
cd \
cd Pentaho/design-tools/data-integration
./spoon.bat
```

Linux:

```powershell
cd \
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Next steps

Browse our comprehensive workshops to learn more about Pentaho solutions.
{% endstep %}
{% endstepper %}

{% columns %}
{% column %}

### Workshops

Pentaho workshops come fully loaded with everything you need to hit the ground running!

Grab ready-to-use solution files, copy-and-paste commands straight into your terminal (no typos, no stress!), and handy automation scripts that do the heavy lifting for you.

<a href="https://academy.pentaho.com/pentaho-data-integration" class="button primary" data-icon="diagram-sankey">Get started with Data Integration</a>&#x20;
{% endcolumn %}

{% column %}
{% code title="index.js" overflow="wrap" %}

```javascript
// Initialize the client
const client = new ExampleAPI({ apiKey: "YOUR_API_KEY" });

// Send your first message
const response = await client.messages.send({
  message: "Hello, world!"
});

```

{% endcode %}
{% endcolumn %}
{% endcolumns %}

<table data-view="cards"><thead><tr><th></th><th align="center"></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-type="image">Cover image (dark)</th></tr></thead><tbody><tr><td><i class="fa-screwdriver-wrench">:screwdriver-wrench:</i>  <a href="https://academy.pentaho.com/installation-of-pentaho"><strong>Installation</strong></a></td><td align="center"></td><td>Install Pentaho Enterprise Edition 11.x on Linux &#x26; Windows.</td><td><a href="/files/j3sQyJl8nNkemtGuejy9">/files/j3sQyJl8nNkemtGuejy9</a></td><td></td></tr><tr><td> <i class="fa-arrow-progress">:arrow-progress:</i>  <a href="https://academy.pentaho.com/pipeline-designer"><strong>Pipeline Designer</strong></a></td><td align="center"></td><td>Pipeline Designer is a web-based interface that  manages data integration pipelines.</td><td><a href="/files/Af2wkqOTpXxswbOfoMKv">/files/Af2wkqOTpXxswbOfoMKv</a></td><td></td></tr><tr><td><i class="fa-diagram-sankey">:diagram-sankey:</i>  <a href="https://academy.pentaho.com/pentaho-data-integration"><strong>Data Integration</strong></a></td><td align="center"></td><td>Simplify hybrid data estates using advanced data pipeline techniques.</td><td><a href="/files/YJ5e3vAD4UrGZsKVlBJm">/files/YJ5e3vAD4UrGZsKVlBJm</a></td><td></td></tr><tr><td> <i class="fa-kaaba">:kaaba:</i>  <a href="https://academy.pentaho.com/semantic-model-editor"><strong>Semantic Model Editor</strong></a></td><td align="center"></td><td>Data Modeling is the first step in understanding the reporting requirements.</td><td><a href="/files/CbN4wwnC8iVnlvouTC26">/files/CbN4wwnC8iVnlvouTC26</a></td><td></td></tr><tr><td> <i class="fa-cube">:cube:</i>  <a href="https://academy.pentaho.com/schema-workbench"><strong>Schema Workbench</strong></a></td><td align="center"></td><td>A visual design tool used to create and edit Mondrian 3 OLAP schemas.</td><td><a href="/files/NCD5ZZrCWh8cCuYpzChC">/files/NCD5ZZrCWh8cCuYpzChC</a></td><td></td></tr><tr><td> <i class="fa-clipboard-list-check">:clipboard-list-check:</i>  <a href="https://academy.pentaho.com/metadata-editor"><strong>Metadata Editor</strong></a></td><td align="center"></td><td>A graphical modeling tool that creates and manages semantic business models.</td><td><a href="/files/P2J4wHyzVwPYg5OZVM6S">/files/P2J4wHyzVwPYg5OZVM6S</a></td><td></td></tr><tr><td><i class="fa-chart-pie-simple">:chart-pie-simple:</i>  <a href="https://academy.pentaho.com/pentaho-business-analytics"><strong>Business Analytics</strong></a></td><td align="center"></td><td>Turn your raw data into actionable insights.</td><td><a href="/files/JAc3lYCBs884RQ4nypaQ">/files/JAc3lYCBs884RQ4nypaQ</a></td><td></td></tr><tr><td><i class="fa-chart-mixed">:chart-mixed:</i>  <a href="https://academy.pentaho.com/pentaho-ctools"><strong>CTools</strong></a></td><td align="center"></td><td>A suite of visualisation tools for creating highly interactive dashboards.</td><td><a href="/files/N5l6lblCLxZ9cdH1x3NX">/files/N5l6lblCLxZ9cdH1x3NX</a></td><td><a href="/files/NaDBPUeMqs499B0dKNM9">/files/NaDBPUeMqs499B0dKNM9</a></td></tr></tbody></table>

{% columns %}
{% column %}
{% embed url="<https://www.loom.com/share/47be2435461240cb82815f173d565514?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" fullWidth="true" %}
Take a tour ..
{% endembed %}
{% endcolumn %}

{% column %}

### Take a Tour

So you've landed on a doc site - let's check out what you can actually do here!&#x20;

Hit Ctrl + K and boom, you've got a search bar that does double duty: type something in for classic keyword search, or click "Ask" to chat with GitBook Assistant ..

{% endcolumn %}
{% endcolumns %}

<h3 align="center">More Resources</h3>

{% columns %}
{% column %}

#### Pentaho Documentation Site

Documenetaion site provides comprehensive technical guides and resources for the Pentaho Platform, covering data integration, analytics, metadata management, data quality, and business intelligence capabilities to help users discover, manage, and derive insights from their data.
{% endcolumn %}

{% column %}

#### Customer Support Portal

Registered customers can:

* download product releases and service packs
* access technical documentation&#x20;
* submit and track support cases
* view security updates
  {% endcolumn %}
  {% endcolumns %}

{% columns %}
{% column %}

<p align="center"><a href="https://docs.pentaho.com/" class="button primary" data-icon="book-open">Documentation</a></p>
{% endcolumn %}

{% column %}

<p align="center"><a href="https://support.pentaho.com/hc/en-us" class="button primary" data-icon="face-smiling-hands">Customer Support Portal</a></p>
{% endcolumn %}
{% endcolumns %}


# Pentaho 11 Installation

Pentaho Enterprise 11 Installation ..

#### Introduction

This page is intended for system administrators, DevOps engineers, and technical users who need to install or evaluate Pentaho Enterprise Edition 11 on Linux or Windows.

{% hint style="info" %}
Pentaho Enterprise Edition 11 is a data integration and analytics platform with tools for data ingestion, transformation, visualization, and reporting. Pentaho Enterprise runs on multiple operating systems, including Linux and Windows.

This page is the starting point for installing Pentaho Enterprise on **Linux** and **Windows**.

For production environments, you can install Pentaho Enterprise in one of the following ways:

* [**Archive (Linux)**](/pentaho-11-installation-en/installation/archive-installation) – Run the Pentaho Server on the Tomcat instance supplied by Pentaho.
* **Manual (advanced)** – Deploy the Pentaho Server on an existing Tomcat or JBoss web application server. This option is intended for experienced administrators and is not covered step-by-step in this workshop.
* [**Client tools**](/pentaho-11-installation-en/installation/archive-installation/install-client-tools) – Install only the Business Analytics (BA) or Data Integration (DI) client components.

For evaluation scenarios on Windows, you can use the guided installer:

* [**Enterprise evaluation (Windows)**](/pentaho-11-installation-en/installation/evaluation-installation) – Install a preconfigured evaluation environment using the Windows wizard.
  {% endhint %}

{% hint style="warning" %}
The Manual installation option is intended for experienced administrators who manage their own application servers (Tomcat or JBoss). This documentation does not provide a full step-by-step guide for Manual installations.
{% endhint %}

***


# Whats New ..

Pentaho 11

Pentaho Data Integration (PDI) 11.0 and Pentaho Business Analytics (PBA) 11.0 deliver major updates. This release focuses on a modern UI, stronger security, and simpler deployments.

11.0 is a long-term support (LTS) release. For service pack and update schedules, see the [Pentaho Support Lifecycle page](https://support.pentaho.com/hc/en-us/articles/205789159-Pentaho-Product-Lifecycle-Overview).

For the full change list, see the [11.0 release notes](https://docs.pentaho.com/whats-new/release-notes-11.0).

### Highlights

* Author ETL pipelines in a browser with Pipeline Designer.
* Organize and deploy ETL assets with project-based lifecycle management.
* Preview the redesigned Pentaho User Console (PUC).
* Build Mondrian models in the browser with Semantic Model Editor (SME).
* Enable SSO with built-in OpenID Connect (OIDC) and OAuth 2.0.
* Simplify container deployments with standardized, prebuilt images.
* Run on Java 21 and Tomcat 10.
* Manage features as plugins with Plugin Manager.
* Monitor ETL runs with OpenTelemetry-based observability.

### Pipeline Designer

Pipeline Designer is a browser-based UI for authoring ETL pipelines. Build PDI transformations and jobs in a web browser. Use Spoon only for legacy tooling.

![Pipeline Designer in the browser](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252FT7qwBRJmSw1J2dkMntsY%252Fimage.png%3Falt%3Dmedia%26token%3De0d86cad-bcd1-40c3-b937-8950f82dc700\&width=768\&dpr=4\&quality=100\&sign=c6becdd6\&sv=2)

It is similar to Spoon. The UI uses a modern framework. It remains compatible with transformations and jobs created in Spoon.

Learn more about [Pipeline Designer](https://docs.pentaho.com/pba/11.0-pba/pipeline-designer).

### Project-based lifecycle management

Before 11.0, PDI had no defined structure for organizing transformations, jobs, and configuration. That made collaboration and environment promotion harder for ETL and DevOps teams.

Configuration resolution could also feel inconsistent.

![Project-based organization for ETL assets](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252Fj9bWqemlvUzaV4Nmz88m%252Fimage.png%3Falt%3Dmedia%26token%3De1e7e58f-70ea-4086-94ea-7a6a70c86332\&width=768\&dpr=4\&quality=100\&sign=d0225709\&sv=2)

Project-based lifecycle management addresses these gaps.

Learn more about [configuring ETL with Projects](https://docs.pentaho.com/pdia-data-integration/pdia-11.0-data-integration/organizing-etl-with-projects#manageability).

### Modern Pentaho User Console (preview)

{% hint style="info" %}
This feature is in preview. Behavior and UI might change in later updates.
{% endhint %}

11.0 introduces a redesigned user experience (UX) for Pentaho User Console (PUC). It aligns with the broader Pentaho platform UX. It addresses pain points in PUC 10.2 and earlier.

![Modern Pentaho User Console](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252FUyUxCG5t0P9X4ealemYT%252Fimage.png%3Falt%3Dmedia%26token%3Df148a13e-3285-46b7-92cf-bfe33e187bd8\&width=768\&dpr=4\&quality=100\&sign=739846af\&sv=2)

The existing PUC stays available until feature parity.

Learn more about [Modern PUC](https://docs.pentaho.com/pba/11.0-pba/pentaho-user-console/modern-design).

### Semantic Model Editor (SME)

11.0 introduces a new tool for building and managing Mondrian data models. Previously, customers used Schema Workbench or Data Source Wizard. Semantic Model Editor (SME) provides a modern, web-based workflow.

SME works for both new and advanced users. It improves the modeling experience in PBA, especially Analyzer. It supports existing Mondrian models and adds new capabilities.

![Semantic Model Editor](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252FbvAaOXY6vTCT8cYvhpLG%252Fimage.png%3Falt%3Dmedia%26token%3Dbf0cdd57-08df-4c7a-8c4a-c2e468dea084\&width=768\&dpr=4\&quality=100\&sign=f676c6f1\&sv=2)

Learn more about [Semantic Model Editor](https://docs.pentaho.com/pba/11.0-pba/semantic-model-editor).

### Built-in OIDC and OAuth 2.0

11.0 supports OpenID Connect (OIDC) and OAuth 2.0 authentication for Pentaho Server. This enables single sign-on (SSO) with identity providers such as Google, Okta, and Azure. It supports any OIDC-compliant identity provider (IdP).

![OIDC/OAuth authentication options](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252FsmMHck3iMCiWXWr3j3jn%252Fimage.png%3Falt%3Dmedia%26token%3Dc66d9337-1950-429e-8847-cc93b64d3850\&width=768\&dpr=4\&quality=100\&sign=7579d4e8\&sv=2)

Learn more about [OIDC and OAuth 2.0](https://docs.pentaho.com/pdia-admin/pdia-11.0-admin/administer/secure-the-pentaho-system/user-security/advanced-security-providers/oidc-oauth-2.0).

### Granular permissions for Pentaho Server

11.0 adds more granular and flexible access control across the platform. This addresses long-standing challenges:

* Permissions were not fine-grained enough. For example, **Read Content** lets users see content from any plugin. File and folder permissions can still block access.
* Permissions could not cleanly allow or block individual plugins.
* Permissions were too broad for data sources and similar assets.
* Execute permissions were too broad.

![Granular permission management](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252FrTbSN2et5tUaVD8nVzls%252Fimage.png%3Falt%3Dmedia%26token%3Dcad4f374-af6b-42d9-8ee9-c9dbd89d865e\&width=768\&dpr=4\&quality=100\&sign=25687fc3\&sv=2)

11.0 addresses these issues in Pentaho Server. Combined with OIDC and OAuth 2.0, it provides a stronger authentication and authorization model.

Learn more about [granular permissions](https://docs.pentaho.com/pba/11.0-pba/semantic-model-editor/sharing-a-semantic-model/permissions-for-semantic-models).

### Simplified container deployment

11.0 simplifies Docker-based deployments. It introduces optimized, prebuilt images for on-premises deployments and major Kubernetes platforms.

This includes plain Docker, Kubernetes, EKS, AKS, and GKE. Images use standardized installation paths and variables. Entrypoint scripts support runtime overrides for configuration files and licenses.

Learn more about [Docker deployment](https://docs.pentaho.com/install/pdia-11.0-installation/pentaho-installation-overview-cp/docker-container-deployment-of-pentaho-installation-cp).

### Java 21 and Tomcat 10 support

Java 21 is supported in 11.0. You can use Oracle JDK, OpenJDK, or other supported JVMs.

Pentaho Server 11.0 ships with Tomcat 10. This addresses vulnerabilities and defects associated with Tomcat 9.

Learn more in the [Components reference](https://docs.pentaho.com/install/pdia-11.0-installation/components-reference).

### Plugin Manager

11.0 introduces a Plugin Manager for both PDI and PBA plugins. Pentaho will ship more functionality as plugins over time. This makes it easier to identify, deploy, and update plugins.

![Plugin Manager](https://docs.pentaho.com/~gitbook/image?url=https%3A%2F%2F2804294592-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FyzYIVXc5NujjLFFhw5Jm%252Fuploads%252Fy5CzkfB31Zn3EuPPsUdn%252Fimage.png%3Falt%3Dmedia%26token%3D8e60d853-c28e-4501-a22d-b5781bc80f17\&width=768\&dpr=4\&quality=100\&sign=6478f59e\&sv=2)

Learn more about [Plugin Manager](https://docs.pentaho.com/pba/11.0-pba/pentaho-user-console/modern-design/plugin-manager).

### Karaf and OSGi removed, big data plugins delivered separately

In 11.0, Karaf and OSGi are removed from PDI. Big data components are now delivered as plugins. There is no separate PDI distribution for big data add-ons. Deploy big data components the same way as other PDI plugins.

This reduces both the PDI client and Pentaho Server deployment size by more than 1 GB.

Learn more about [Plugin Manager](https://docs.pentaho.com/pba/11.0-pba/pentaho-user-console/modern-design/plugin-manager).

### OpenTelemetry-based observability

[OpenTelemetry](https://opentelemetry.io/docs/languages/java/) (OTel) is an open standard for sharing telemetry data such as traces, metrics, and logs. Many tools can consume OpenTelemetry data, including Datadog, Splunk, Elastic, Amazon CloudWatch, and Azure Monitor.

With the OTel plugin, you can monitor Pentaho ETL processes with:

* Logs that are consolidated in a single place and represented hierarchically
* Traces to view task timing, execution hierarchy, and variables during execution
* Metrics to track data flow trends at specified points of interest

Learn more about [Plugin Manager](https://docs.pentaho.com/pba/11.0-pba/pentaho-user-console/modern-design/plugin-manager).

### Other enhancements

11.0 also includes smaller enhancements and defect fixes. See the [11.0 release notes](https://docs.pentaho.com/whats-new/release-notes-11.0) for details.


# Getting Started

Archive installation of Pentaho Enterprise on Linux - Ubuntu 24.04 LTS ..

{% stepper %}
{% step %}
**Download packages (EE)**

Select your download binaries option:

{% tabs %}
{% tab title="Enterprise Edition" %}
{% hint style="info" %}
**30-day Enterprise Edition**

The 30-day Pentaho Enterprise download is for:
{% endhint %}

1. Navigate to the Pentaho download page & fill in the form.

{% embed url="<https://pentaho.com/download/>" %}

<figure><img src="/files/oAHp1iMCehAii6a4LBAl" alt=""><figcaption><p>Fill out form</p></figcaption></figure>

2. Download Pentaho EE onPrem

<figure><img src="/files/6vjl1tbOQIbGDYENotW4" alt=""><figcaption><p>Download - Pentaho EE onPrem</p></figcaption></figure>

3. Select the version for your OS:

<figure><img src="/files/gr2cfedErWWTGvY59MsD" alt=""><figcaption><p>Pentaho Enterprise OS versions</p></figcaption></figure>

{% hint style="info" %}
You can download more than OS version.
{% endhint %}
{% endtab %}

{% tab title="Pentaho Support Portal" %}
{% hint style="info" %}
**Pentaho Customer Portal (EE)**

Access Pentaho Enterprise downloads with your principal account credentials.
{% endhint %}

1. Log into Pentaho Customer Portal.

{% embed url="<https://support.pentaho.com/hc/en-us>" %}

<figure><img src="/files/iSsMHdiuwlZFyP8FOS2t" alt=""><figcaption><p>Pentaho Support Portal</p></figcaption></figure>

2. Click Pentaho > DOWNLOAD.
3. Sign In to Pentaho Customer Portal.

<figure><img src="/files/ZVNXnn4walx3sUXfDd2i" alt=""><figcaption><p>Pentaho Customer Support Portal - Log In</p></figcaption></figure>

4. Select your Pentaho version: Pentaho 11.0 GA Release

<figure><img src="/files/LQv73QeQgeY156Hx7Dn7" alt=""><figcaption></figcaption></figure>

5. Select the required Pentaho packages to download.

<figure><img src="/files/c75KafNayncc4L4ZfDJy" alt=""><figcaption><p>Packages</p></figcaption></figure>

{% hint style="info" %}
**Pentaho packages used in this lab:**

**Big Data Shims**

* `pentaho-server-ee-11.0.0.0-237.zip`

**Pentaho Server**

* `pentaho-server-ee-11.0.0.0-237.zip`
* `paz-plugin-ee-11.0.0.0-237.zip`
* `pir-plugin-ee-11.0.0.0-237.zip`
* `pdd-plugin-ee-11.0.0.0-237.zip`
* `webttle-plugin-ee-11.0.0.0-237.zip`
* `semantic-model-editor-plugin-assembly-1.0.0.zip`
* `pas-scheduler-11.0.0.0-237.zip`
* `webttle-carte-api-plugin-11.0.0.0-237.zip`

**Client Tools**

* `pdi-ee-11.0.0.0-237.zip`
* `pad-ee-11.0.0.0-237.zip`
* `psw-ee-11.0.0.0-237.zip`
* `pme-ee-11.0.0.0-237.zip`
* `prd-ee-11.0.0.0-237.zip`

**Operations Mart**

* `pentaho-operations-mart-11.0.0.0-237.zip`

**Docker Image Configurator**

* `on-prem-11.0.0.0-237.zip`

Note: Additional EE plugins, if required for your use case, will be installed later in [Server Plugins](/pentaho-11-installation-en/installation/archive-installation/install-pentaho-server/server-plugins).
{% endhint %}
{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}
**License Key Information**

* Your License Key is already activated with the 30-day Pentaho Enterprise Edition . This License Key is applicable for any of the supported OS.
* For **Pentaho Partners only**, request a license from your regional Pentaho Partner Manager.
  {% endstep %}

{% step %}
**Install Pentaho Enterprise Edition**

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="image">Cover image</th></tr></thead><tbody><tr><td><a href="https://academy.pentaho.com/installation-of-pentaho-11/installation/evaluation-installation">Windows 11 Installation</a></td><td><a href="/files/45hzmlZbZPOvPHCR7ZIV">/files/45hzmlZbZPOvPHCR7ZIV</a></td></tr><tr><td></td><td></td></tr><tr><td><a href="/pages/cTu381Tu5j7GsbWNNn90">Linux Archive Installation</a></td><td data-object-fit="fill"><a href="/files/H9P2DRCqPfgYBRFbNkGY">/files/H9P2DRCqPfgYBRFbNkGY</a></td></tr></tbody></table>
{% endstep %}

{% step %}
**Try some of the Pentaho Workshops**

Pentaho workshops come fully loaded with everything you need to hit the ground running!

Grab ready-to-use solution files, copy-and-paste commands straight into your terminal (no typos, no stress!), and handy automation scripts that do the heavy lifting for you.

<a href="https://academy.pentaho.com/pentaho-data-integration" class="button primary" data-icon="diagram-sankey">Get started with Data Integration</a>

<table data-view="cards"><thead><tr><th></th><th align="center"></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-type="image">Cover image (dark)</th></tr></thead><tbody><tr><td><i class="fa-screwdriver-wrench">:screwdriver-wrench:</i> <a href="https://academy.pentaho.com/installation-of-pentaho"><strong>Installation</strong></a></td><td align="center"><mark style="color:$danger;"><strong>NEW</strong></mark></td><td>Install Pentaho Enterprise Edition 11.x on Linux &#x26; Windows.</td><td><a href="/files/ieWw0fkisLJYDec1ZxlZ">/files/ieWw0fkisLJYDec1ZxlZ</a></td><td></td></tr><tr><td><i class="fa-arrow-progress">:arrow-progress:</i> <a href="https://academy.pentaho.com/pipeline-designer"><strong>Pipeline Designer</strong></a></td><td align="center"><mark style="color:$danger;"><strong>NEW</strong></mark></td><td>Pipeline Designer is a web-based interface that manages data integration pipelines.</td><td><a href="/files/Af2wkqOTpXxswbOfoMKv">/files/Af2wkqOTpXxswbOfoMKv</a></td><td></td></tr><tr><td><i class="fa-diagram-sankey">:diagram-sankey:</i> <a href="https://academy.pentaho.com/pentaho-data-integration"><strong>Data Integration</strong></a></td><td align="center"></td><td>Simplify hybrid data estates using advanced data pipeline techniques.</td><td><a href="/files/YJ5e3vAD4UrGZsKVlBJm">/files/YJ5e3vAD4UrGZsKVlBJm</a></td><td></td></tr><tr><td><i class="fa-kaaba">:kaaba:</i> <a href="https://academy.pentaho.com/semantic-model-editor"><strong>Semantic Model Editor</strong></a></td><td align="center"><mark style="color:$danger;"><strong>NEW</strong></mark></td><td>Data Modeling is the first step in understanding the reporting requirements.</td><td><a href="/files/CbN4wwnC8iVnlvouTC26">/files/CbN4wwnC8iVnlvouTC26</a></td><td></td></tr><tr><td><i class="fa-cube">:cube:</i> <a href="https://academy.pentaho.com/schema-workbench"><strong>Schema Workbench</strong></a></td><td align="center"></td><td>A visual design tool used to create and edit Mondrian 3 OLAP schemas.</td><td><a href="/files/NCD5ZZrCWh8cCuYpzChC">/files/NCD5ZZrCWh8cCuYpzChC</a></td><td></td></tr><tr><td><i class="fa-clipboard-list-check">:clipboard-list-check:</i> <a href="https://academy.pentaho.com/metadata-editor"><strong>Metadata Editor</strong></a></td><td align="center"></td><td>A graphical modeling tool that creates and manages semantic business models.</td><td><a href="/files/P2J4wHyzVwPYg5OZVM6S">/files/P2J4wHyzVwPYg5OZVM6S</a></td><td></td></tr><tr><td><i class="fa-chart-pie-simple">:chart-pie-simple:</i> <a href="https://academy.pentaho.com/pentaho-business-analytics"><strong>Business Analytics</strong></a></td><td align="center"></td><td>Turn your raw data into actionable insights.</td><td><a href="/files/JAc3lYCBs884RQ4nypaQ">/files/JAc3lYCBs884RQ4nypaQ</a></td><td></td></tr><tr><td><i class="fa-chart-mixed">:chart-mixed:</i> <a href="https://academy.pentaho.com/pentaho-ctools"><strong>CTools</strong></a></td><td align="center"></td><td>A suite of visualisation tools for creating highly interactive dashboards.</td><td><a href="/files/N5l6lblCLxZ9cdH1x3NX">/files/N5l6lblCLxZ9cdH1x3NX</a></td><td><a href="/files/JAc3lYCBs884RQ4nypaQ">/files/JAc3lYCBs884RQ4nypaQ</a></td></tr></tbody></table>

{% embed url="<https://www.loom.com/share/47be2435461240cb82815f173d565514?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" fullWidth="true" %}
Take a tour ..
{% endembed %}
{% endstep %}
{% endstepper %}

***


# Components

Overview of Pentaho Enterprise components ..

{% hint style="info" %}

#### **Pentaho Client / Server Architecture**

Pentaho's client/server architecture forms the basis of its data integration and business analytics suite, providing a flexible and scalable platform for enterprise data management and analysis. The architecture is designed to support various data integration, reporting, and analytics needs across an organization.
{% endhint %}

<figure><img src="/files/HiKwECEzwO4g2cxBfMQd" alt=""><figcaption><p>Pentaho Architecture</p></figcaption></figure>

<table><thead><tr><th width="186">Port Number</th><th>Description</th></tr></thead><tbody><tr><td>5432</td><td>PostgreSQL Server</td></tr><tr><td>8080</td><td>Pentaho Server Tomcat Web Server Startup Port</td></tr><tr><td>8012</td><td>Pentaho Server Shutdown Port</td></tr><tr><td>9001</td><td>HSQL Server Port</td></tr><tr><td>9092</td><td>Embedded H2 Database</td></tr></tbody></table>

Key components include:

{% tabs %}
{% tab title="Pentaho Client Tools" %}
{% hint style="info" %}

#### **Pentaho Client Tools**

The Pentaho Client, a key component of the Pentaho suite, encompasses several user-facing tools designed for data management and analytics. These include the Data Integration tool (PDI), which is central to extracting, transforming, and loading (ETL) operations; Spoon, a graphical user interface for designing ETL processes; Designer for convenient pipeline design; Scheduler linked to Quartz for job scheduling; Repository Browser for managing ETL assets; and Database Explorer for database operations.

Additionally, it offers tools like Metadata Editor and Schema Workbench for advanced data manipulation. Together, these tools empower users to efficiently process and analyze data within the Pentaho ecosystem.
{% endhint %}

{% tabs %}
{% tab title="Data Integration" %}
{% hint style="info" %}

#### **Data Integration**

Pentaho Data Integration (PDI), also known as Kettle, is an open-source data integration tool that allows the extraction, transformation, and loading (ETL) of data into databases, data warehouses, and business applications. It is designed to handle a wide variety of data sources including traditional relational databases, unstructured data formats, and cloud-based storage. PDI is composed of several key components that work together to provide a comprehensive ETL solution.
{% endhint %}

<figure><img src="/files/cXzgGRAak6e2aiO0OeF3" alt=""><figcaption><p>Pentaho Client / Server Architecture</p></figcaption></figure>

{% hint style="info" %}
**Spoon**

Spoon is the graphical user interface (GUI) for designing and testing PDI jobs and transformations. It allows users to visually create, edit, and manage ETL processes without writing code.

**Designer**

Drag & Drop 'objects' to design your pipelines and workflows.

**Scheduler**

Connects to Quartz scheduler on server. Jobs and transformations must be uploaded to Repository.

**Repository Browser**

The repository is a central storage area for PDI resources such as jobs, transformations, and database connections. It facilitates collaboration among team members by allowing them to share and manage ETL assets efficiently.

These components collectively make PDI a powerful tool for data integration, enabling businesses to cleanse, integrate, and analyze data from diverse sources more effectively.

Connects to Apache Jackrabbit content Repository, pointing to a supported database:

* PostgreSQL
* MSSQL Server
* Oracle
* MySQL
* MariaDB

**DB Explorer**

Database Explorer that enables you to conduct minimal database operations.
{% endhint %}
{% endtab %}

{% tab title="Metadata Editor" %}
{% hint style="info" %}

#### **Metadata Editor**

The Pentaho Metadata Editor is a tool within the Pentaho suite that facilitates the creation and management of business models. These models form the foundation for reporting and analysis, making it easier for end-users to interact with data without needing a deep understanding of the underlying database structures.

Key features include:

**User-friendly Interface:** Offers a graphical environment where users can define business models, relationships, and metadata concepts, simplifying complex data structures into more understandable terms.

**Data Source Connection:** Allows connection to various data sources, enabling the extraction of metadata from relational databases, OLAP sources, and more.

**Security Settings:** Supports the definition of security constraints at the model level, ensuring that sensitive data remains protected and access is controlled.

**Localization and Internationalization:** Models can be localized, allowing the presentation of metadata in different languages to support global deployments.

The Metadata Editor plays a crucial role in the Pentaho Business Analytics suite, streamlining the creation of complex reports and analyses by offering a simplified view of data for business users.
{% endhint %}

<figure><img src="/files/8JLiUd7EisnUm5iJ7zHK" alt=""><figcaption><p>Metadata Editor</p></figcaption></figure>
{% endtab %}

{% tab title="Schema Workbench" %}
{% hint style="info" %}

#### **Schema Workbench**

The Pentaho Schema Workbench is an essential tool within the Pentaho suite designed for developers and data architects to create and edit OLAP (Online Analytical Processing) schemas. It provides a graphical interface for defining the multidimensional models needed for complex analytical queries, enabling the efficient organization and visualization of large data sets.

With its user-friendly interface, users can easily design OLAP cubes that form the foundation of advanced analytics and business intelligence applications, making data more actionable and insights more accessible.
{% endhint %}

<figure><img src="/files/L65jQf0L4Ql7onVNWZOR" alt=""><figcaption><p>Schema Workbench</p></figcaption></figure>
{% endtab %}

{% tab title="Aggregation Designer" %}
{% hint style="info" %}

#### **Aggregation Designer**

The Pentaho Aggregation Designer is a pivotal tool aimed at improving query performance by simplifying the creation and management of aggregate tables in a star schema database. This graphical tool assists users in defining, generating, and deploying SQL-based aggregation tables that summarily condense detailed data into summarized formats, making data retrieval processes significantly more efficient for analytical queries.

This capability is critical for enhancing the performance of OLAP cubes, facilitating faster data analysis, and providing a more streamlined user experience in the Pentaho Business Analytics suite.
{% endhint %}

<figure><img src="/files/gDsvcQLdlLbBCKeL2qgx" alt=""><figcaption><p>Aggregation Designer</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Pentaho Server" %}
{% hint style="info" %}
**Pentaho Server**

Pentaho Server acts as the central platform for hosting and managing all Pentaho applications and services. It provides a secure, scalable environment for deploying and executing Pentaho's analytics and data integration solutions. Key components include:

* **BI Server:** Facilitates interactive reporting, analytics, dashboarding, and data exploration.
* **Data Integration Server:** Supports the orchestration and scheduling of ETL (Extract, Transform, Load) processes.
* **User Console:** Offers a web-based interface for accessing, creating, and managing content within the Pentaho suite.
* **Security:** Integrates with enterprise security systems to provide authentication, authorization, and secure access.
* **Repository:** Centralizes the storage of all Pentaho assets, including reports, dashboards, and ETL scripts, ensuring collaboration and version control.

The server enables organizations to leverage the full potential of the Pentaho suite by providing a comprehensive platform for business intelligence and data management activities.
{% endhint %}

**Pentaho Server Reporting Suite**

{% tabs %}
{% tab title="Analyzer" %}
{% hint style="info" %}

#### **Analyzer**

Pentaho Analyzer is an interactive analytics and data visualization tool that is part of the Pentaho Business Analytics suite. It enables users to explore and analyze data through an intuitive web-based interface, providing rich graphical representations of data including charts, tables, and heat maps. Users can create and customize reports and dashboards without the need for in-depth technical knowledge, making it accessible to a wide range of users. Key features include:

* **Ad-hoc analysis:** Empowers users to quickly create and modify reports based on their specific questions and needs.
* **Drag-and-drop interface:** Simplifies the process of designing reports by allowing users to easily select and arrange data elements.
* **Rich visualizations:** Supports a wide array of visualization options to help users uncover insights from their data.
* **Collaboration and sharing:** Enables sharing of reports and dashboards with other users to facilitate decision-making across teams and departments.

Pentaho Analyzer is designed to work seamlessly with the Pentaho suite, integrating directly with Pentaho's data integration, ETL, and data warehousing capabilities. This allows users to leverage the full power of the suite for comprehensive data analysis and business intelligence solutions.
{% endhint %}

<figure><img src="/files/pJb1oZOCGiQHgrZD1bV9" alt=""><figcaption><p>Analyzer Report</p></figcaption></figure>
{% endtab %}

{% tab title="Interactive Reports" %}
{% hint style="info" %}

#### **Interactive Reports**

Pentaho Interactive Reports offer a highly user-friendly interface for creating, editing, and viewing ad-hoc reports. This feature is designed for business users who need to generate reports quickly without in-depth technical knowledge of the underlying data structure.

* **User-Friendly Interface:** Provides a drag-and-drop interface, making it easy for users to select, organize, and present data without any SQL knowledge.
* **Real-Time Data Exploration:** Enables users to interact with their data in real-time, allowing for instant filtering, sorting, and aggregation to identify trends and insights.
* **Customizable Layouts:** Users can customize the layout of their reports by adjusting columns, rows, and summaries to meet their specific reporting needs.
* **Export and Share:** Reports can be exported to various formats (e.g., PDF, Excel, CSV) and shared with stakeholders to support data-driven decision-making.

Interactive Reports are part of the larger Pentaho Business Analytics suite, offering seamless integration with Pentaho's ETL and data analysis tools, ensuring businesses have a comprehensive solution for their data integration and reporting needs.
{% endhint %}

<figure><img src="/files/Yel0hZU1S2OMDsONIwEp" alt=""><figcaption><p>Interactive Report</p></figcaption></figure>
{% endtab %}

{% tab title="Dashboard Designer" %}
{% hint style="info" %}

#### **Dashboard Designer**

Pentaho Dashboard Designer is a feature-rich tool within the Pentaho Business Analytics suite, designed for creating interactive and visually appealing dashboards. These dashboards aggregate and display data from various sources, providing users with insights at a glance. Here's a quick overview:

* **Intuitive Design Interface**: Offers a drag-and-drop interface, making it accessible for non-technical users to create and customize dashboards.
* **Data Integration**: Seamlessly integrates with Pentaho Data Integration (PDI), allowing it to pull data from a wide range of sources for real-time analytics.
* **Interactive Widgets**: Supports various types of widgets including charts, tables, and filters, enabling interactive data exploration.
* **Customization and Branding**: Allows for the customization of layout and design, enabling alignment with company branding.
* **Collaboration Features**: Facilitates sharing and collaboration by allowing users to publish dashboards within the organization or to a broader audience.
* **Security**: Integrates with existing security frameworks, ensuring data protection and controlled access based on roles and permissions.

Pentaho Dashboard Designer plays a crucial role in transforming data into actionable insights, driving informed decision-making across organizations.
{% endhint %}

<figure><img src="/files/OO1bxDc6u4E5PPy4b3nZ" alt=""><figcaption><p>Dashboard</p></figcaption></figure>
{% endtab %}

{% tab title="Pipeline Designer" %}
{% hint style="info" %}

#### Pipeline Designer

{% endhint %}

<figure><img src="/files/NoPrRLZjjmDoYUBGBWhZ" alt=""><figcaption><p>Pipeline Designer</p></figcaption></figure>
{% endtab %}

{% tab title="Semantic Model Editor" %}
{% hint style="info" %}

#### Semantic Model Editor

{% endhint %}

x
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Carte Server" %}
{% hint style="info" %}
**Carte Server**

Pentaho Carte is a lightweight web server for remote execution and monitoring of ETL processes created in Pentaho Data Integration (PDI/Kettle).

Carte is built on Java and uses the embedded Jetty web server. It relies on XML-based configuration and exposes functionality through a REST API, with a simple browser-based interface for monitoring.

The server enables remote execution of transformations and jobs, supports clustering for load balancing, provides real-time monitoring, and allows scheduling of ETL processes.

Carte can be deployed as a standalone server, in a master-slave cluster setup, or in a load-balanced environment for high availability. It's typically launched via command line with a configuration file containing server settings.

This component is crucial for Pentaho's distributed processing architecture, allowing organizations to scale data integration processes across multiple machines.
{% endhint %}

<figure><img src="/files/AlkihPLbR8jLYrbnFo06" alt=""><figcaption><p>Carte Cluster</p></figcaption></figure>
{% endtab %}

{% tab title="APIs" %}
{% hint style="info" %}
**Kitchen**

Kitchen is a command-line tool that enables the execution of PDI jobs. It supports batch processing and can be integrated into automated workflows, allowing for efficient data processing.
{% endhint %}

```
kitchen.sh -file=/PRD/updateWarehouse.kjb -level=Minimal
kitchen.bat /file:D:\Jobs\updateWarehouse.kjb /level:Basic
```

{% hint style="info" %}
**Pan**

Similar to Kitchen, Pan is a command-line tool but is specifically designed for executing PDI transformations. It provides flexibility in running ETL transformations from shell scripts or scheduling systems.
{% endhint %}

```
pan.sh -file="/PRD/Customer Dimension.ktr" -level=Minimal
pan.bat /file:"D:\Transformations\Customer Dimension.ktr" /level:Basic
```

{% embed url="<https://docs.pentaho.com/pentaho-rest-api>" %}
{% endtab %}
{% endtabs %}


# Archive Installation

Recommended installation option ..

#### Overview

Choose the Archive installation when you:

* Don’t already run Tomcat and want an opinionated, faster setup.
* Need a straightforward path to migrate content from an existing Pentaho environment.

{% hint style="info" %}
Archive installation ships Pentaho v11 on Tomcat 10 as a snapshot, speeding setup versus a manual app‑server install.
{% endhint %}

#### What you’ll achieve

* Pentaho Server running on Tomcat 10
* A configured Pentaho Repository (database) - PostgreSQL 17.7
* Server plugins installed
* Client tools installed
* Licenses applied and validated

#### Who is this for?

* Admins who want a quick, supported Tomcat + Pentaho setup
* Teams migrating content from previous Pentaho versions or evaluation installs

#### Prerequisites

* Ubuntu 24.04 LTS with sudo access
* Java 21 (OpenJDK) installed and `PENTAHO_JAVA_HOME` set
* Access to a supported database for the Pentaho Repository (PostgreSQL recommended)
* Verify supported versions: [Components Reference](https://docs.pentaho.com/install/pdia-11.0-installation/components-reference)

{% hint style="danger" %}
Uninstall any evaluation versions of Pentaho before proceeding.
{% endhint %}

{% stepper %}
{% step %}
**Prepare your environment**

* Create a dedicated installation user `pentaho` and directory layout.
* Install and validate Java 21; set `PENTAHO_JAVA_HOME`.
* Install and configure your Pentaho Repository database (PostgreSQL recommended).

Go to: [Prepare Environment](/pentaho-11-installation-en/installation/archive-installation/prepare-environment)
{% endstep %}

{% step %}
**Install Pentaho Server (Archive)**

* Download and unpack the archive under `/opt/pentaho`.
* Configure Repository connectivity and Tomcat.
* Start the server and verify logs.

Go to: [Install Pentaho Server](/pentaho-11-installation-en/installation/archive-installation/install-pentaho-server)
{% endstep %}

{% step %}
**Install server plugins**

* Add reporting/visualization plugins required by your use cases.
* Add Semantic Model Editor (SME) for data modeling.
* Add Pipeline Designer (PPD), Scheduler and Carte for creating and deploying automated data pipelines.
* Restart and validate.

Go to: [Server Plugins](/pentaho-11-installation-en/installation/archive-installation/install-pentaho-server/server-plugins)
{% endstep %}

{% step %}
**Install client tools**

* Install PDI, PRD, PME, PSW or other client tools used by your team.

Go to: [Install Client Tools](/pentaho-11-installation-en/installation/archive-installation/install-client-tools)
{% endstep %}

{% step %}
**Start the Pentaho Server and apply licenses**

* Start Pentaho, access PUC, and validate basic functionality.
* Apply licenses via the License Manager.

See: “Start Server” and “License Manager” in [Install Pentaho Server](/pentaho-11-installation-en/installation/archive-installation/install-pentaho-server)
{% endstep %}

{% step %}
**Post‑installation hardening (recommended)**

* Secure credentials, restrict access, and tune performance for production.

Go to: [Post Installation Tasks](/pentaho-11-installation-en/installation/post-installation-tasks)
{% endstep %}
{% endstepper %}

{% hint style="info" %}
Migrating content? Plan your repository upgrade/restore and test before opening the system to end users.
{% endhint %}

***


# Prepare Environment

Preflight tasks ..

{% hint style="info" %}

#### **Prepare Environment**

Prepare your Ubuntu server for an Archive installation of the Pentaho Server.

This process will:

* Create a `pentaho` installation user (with sudo)
* Set Pentaho paths - `$PENTAHO_BASE`
* Install Java 21 (OpenJDK)
* Set `PENTAHO_JAVA_HOME`
* Install PostgreSQL 17 (15 supported for Pentaho 11)
* Create repository database user - `pentaho` for setup
* Install pgAdmin 4 (desktop)
  {% endhint %}

{% hint style="warning" %}
Supported Linux baseline for Pentaho 11.x: Ubuntu 24.04 LTS.

For versions and compatibility, see [Components Reference](https://docs.pentaho.com/install/pdia-11.0-installation/components-reference).
{% endhint %}

{% stepper %}
{% step %}
**Prerequisites**

1. Ensure unzip is installed.

```bash
unzip --version
```

2. Set the Pentaho path variables.

{% hint style="info" %}
Use these variables to simplify commands and avoid path mistakes. Consider adding them to your shell profile for convenience.
{% endhint %}

* Edit the /etc/environment file.

```bash
cd
cd /etc/
sudo nano environment
```

* Add the following paths.

```bash
# Pentaho paths
PENTAHO_BASE=/opt/pentaho
PENTAHO_SERVER=/opt/pentaho/server/pentaho-server
TOMCAT_HOME=/opt/pentaho/server/pentaho-server/tomcat
```

{% hint style="warning" %}
If you add these to `~/.bashrc` or `/etc/environment`, re-open your shell or `source` the file to apply.
{% endhint %}

<figure><img src="/files/VJWdS6gMX3tCkW0UVDVP" alt=""><figcaption><p>Pentaho paths</p></figcaption></figure>
{% endstep %}

{% step %}
**Create a Pentaho installation user**

{% hint style="danger" %}
For production, use a dedicated installation account with only the required privileges and avoid sharing credentials.
{% endhint %}

1. Update packages (run once):

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Add the user and set a password:

```bash
sudo adduser pentaho
```

3. Grant sudo privileges:

```bash
sudo usermod -aG sudo pentaho
```

4. Validate access:

```bash
su - pentaho
groups
sudo -v
```

{% endstep %}

{% step %}
**Install Java 21 (OpenJDK)**

{% hint style="info" %}
Pentaho 11.x is certified with Java 21.
{% endhint %}

<details>

<summary>What's the difference bewteen Oracle JDK &#x26; OpenJDK?</summary>

Oracle JDK and OpenJDK are both implementations of the Java Platform, but they have some important differences:

**Licensing and Cost:**

* **OpenJDK** is completely free and open source under the GPL license. You can use it for any purpose without restrictions.
* **Oracle JDK** changed its licensing model in 2019. It's now free for development and personal use, but requires a paid subscription for commercial production use (Oracle Java SE Subscription).

**Source and Development:**

* **OpenJDK** is the reference implementation of Java and serves as the base for most JDK distributions. Oracle actually contributes significantly to OpenJDK development.
* **Oracle JDK** is built from the OpenJDK source code but includes some additional proprietary components and commercial features.

**Performance and Features:**

* In modern versions (Java 11+), the performance differences are negligible. Oracle has contributed most of its performance improvements back to OpenJDK.
* Oracle JDK historically included some additional tools and features (like Java Flight Recorder and Java Mission Control), but many of these have been open-sourced and are now available in OpenJDK.

**Support and Updates:**

* **OpenJDK** receives community support and updates for about 6 months per release (except for LTS versions maintained by various vendors).
* **Oracle JDK** offers Long Term Support (LTS) with commercial subscriptions, providing updates and security patches for extended periods.

**Other Distributions:** Many vendors offer their own builds of OpenJDK with long-term support, including Amazon Corretto, Azul Zulu, Eclipse Temurin (formerly AdoptOpenJDK), and Red Hat OpenJDK.

For most developers and organizations, OpenJDK or vendor-supported OpenJDK distributions are the go-to choice unless you specifically need Oracle's commercial support.

</details>

1. Install Java 21:

```bash
sudo apt install -y openjdk-21-jre-headless
```

2. Verify Java:

```bash
java -version
which java
readlink -f $(which java)
```

<figure><img src="/files/8AQZcweA5zjX5c1ztTt9" alt="Ubuntu apt installed OpenJDK versions list"><figcaption><p>OpenJDK versions</p></figcaption></figure>

<figure><img src="/files/atDVLvl9uMDLXvW2n0HM" alt="java -version output showing Java 21"><figcaption><p>Java 21</p></figcaption></figure>

{% hint style="info" %}
If multiple Java versions are installed, select the default:

```bash
sudo update-alternatives --config java
```

{% endhint %}
{% endstep %}

{% step %}
**Set `PENTAHO_JAVA_HOME`**

Set `PENTAHO_JAVA_HOME` globally so the Pentaho Server consistently uses Java 21.

1. Edit `/etc/environment`:

```bash
sudo nano /etc/environment
```

2. Add (or update) the following line:

```bash
# set PENTAHO_JAVA_HOME
PENTAHO_JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64
```

<figure><img src="/files/ybe9eCLgooJyXDVGp0XG" alt=""><figcaption><p>Set PENTAHO_JAVA_HOME</p></figcaption></figure>

3. Save and reload the environment (or log out/in):

```bash
source /etc/environment
```

4. Verify:

```bash
echo $PENTAHO_JAVA_HOME
```

***

{% hint style="info" %}
Alternative: set for a single user only in `~/.bashrc`.
{% endhint %}

1. Edit:

```bash
nano ~/.bashrc
```

2. Append and apply:

```bash
export PENTAHO_JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64
. ~/.bashrc
```

{% endstep %}

{% step %}
**Install PostgreSQL 17**

{% hint style="info" %}
Ubuntu 24.04’s default repository provides a newer PostgreSQL 16. To install 17, add the official PostgreSQL (PGDG) repository.

If a different PostgreSQL is already present, purge it first to avoid port and package conflicts (see the optional "Clean previous installs" tab).

**Prerequisites**\
Ubuntu 24.04\
Root privileges or sudo access\
use `sudo su` to get into root instead of pentaho (default user)
{% endhint %}

{% tabs %}
{% tab title="Install PostgreSQL 17" %}

1. Before installing PostgreSQL, ensure your system is up to date.

```bash
sudo apt update && sudo apt upgrade -y
```

2. Install prerequisite packages.

```bash
sudo apt install -y wget ca-certificates
```

3. Import PostgreSQL GPG Key.

```bash
wget --quiet -O - https://www.postgresql.org/media/keys/ACCC4CF8.asc | sudo apt-key add -
```

4. Add PostgreSQL Repository

```bash
sudo sh -c 'echo "deb http://apt.postgresql.org/pub/repos/apt $(lsb_release -cs)-pgdg main" > /etc/apt/sources.list.d/pgdg.list'
```

5. Update Package List.

```bash
sudo apt update
```

6. Install PostgreSQL 17 server and client packages.

```bash
sudo apt update -y && sudo apt upgrade -y
sudo apt install -y postgresql-17 postgresql-contrib-17
```

{% hint style="info" %}
**What gets installed:**

* `postgresql-17`: Main database server
* `postgresql-contrib-17`: Additional utilities and extensions
  {% endhint %}

7. Check service and version:

```bash
sudo systemctl status postgresql --no-pager
psql --version
```

<figure><img src="/files/EK8OpfHzsuIGsIIbix25" alt=""><figcaption><p>PostgreSQL 17.7</p></figcaption></figure>

8. Optional tidy up:

```bash
sudo apt autoremove -y
```

{% endtab %}

{% tab title="Optional: Install PostgreSQL 16" %}

1. Ensure your Ubuntu system is up-to-date.

```bash
sudo apt update && sudo apt upgrade
```

2. To assist in installing the database software, install the following packages.

```bash
sudo apt install dirmngr ca-certificates software-properties-common apt-transport-https lsb-release curl -y
```

3. Import the repository signing key.

```bash
sudo apt install curl ca-certificates gnupg
curl https://www.postgresql.org/media/keys/ACCC4CF8.asc | sudo gpg --dearmor | sudo tee /etc/apt/trusted.gpg.d/apt.postgresql.org.gpg >/dev/null
```

4. Create the repository configuration file.

```bash
echo "deb [signed-by=/etc/apt/trusted.gpg.d/apt.postgresql.org.gpg] http://apt.postgresql.org/pub/repos/apt/ noble-pgdg main" | sudo tee /etc/apt/sources.list.d/pgdg.list
```

5. Update your package list again to include the new repository.

```bash
sudo apt update
```

6. Now, install the PostgreSQL 16 server package.

```bash
sudo apt install -y postgresql-16 postgresql-client-16
psql --version
```

{% endtab %}

{% tab title="Optional: Clean previous installs" %}
If an unsupported PostgreSQL is installed, purge it first. Run each command separately and review output:

```bash
sudo apt-get --purge remove postgresql
sudo apt-get purge postgresql*
sudo apt-get --purge remove postgresql postgresql-doc postgresql-common
sudo apt autoremove -y
```

Check residual packages:

```bash
dpkg -l | grep -i postgres
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
Validation:

```bash
sudo systemctl is-enabled postgresql
sudo ss -ltnp | grep 5432 || true
sudo -u postgres psql -c "SELECT version();"
```

{% endhint %}
{% endstep %}

{% step %}
**Create database user - pentaho**

{% hint style="info" %}
During install, PostgreSQL creates a local superuser `postgres`. Set its password and (optionally) creates a dedicated `pentaho` role.
{% endhint %}

{% hint style="danger" %}
For production, avoid a permanent superuser for application use. Use least-privilege roles and grant only what is needed. The workshop path below elevates privileges to simplify setup.
{% endhint %}

1. Switch to the postgres system user.

```bash
sudo -i -u postgres
```

2. Enter the PostgreSQL interactive terminal.

```bash
psql
```

3. You should see the PostgreSQL prompt:

```
postgres=#
```

4. View current databases.

```
\l
```

{% hint style="info" %}
This lists all databases. You should see three default databases: postgres, template0, and template1.
{% endhint %}

5. Exit.

```sql
q
```

6. Check PostgreSQL Version from SQL.

```sql
SELECT version();
```

7. Exit PostgreSQL Prompt.

<pre class="language-sql"><code class="lang-sql"><strong>\q
</strong></code></pre>

8. And Exit psql.

```sql
exit
```

***

**Create pentaho user**

<figure><img src="/files/KBFAMYiriAeVOtYFcmWs" alt=""><figcaption></figcaption></figure>

1. Set `postgres` password:

```bash
sudo -u postgres psql -c "ALTER USER postgres WITH PASSWORD 'SecurePassword123';"
```

2. Create a `pentaho` user and grant `SUPERUSER` privileges:

```sql
sudo -u postgres psql -c "CREATE USER pentaho WITH PASSWORD 'SecurePassword123';"
sudo -u postgres psql -c "ALTER USER pentaho WITH SUPERUSER;" # Demo only
```

{% hint style="info" %}

* `sudo -u postgres` - Run the command as the Linux system user `postgres` (who has local access to PostgreSQL without a password)
* `psql -c` - Execute a single SQL command and exit
* `CREATE USER pentaho` - Creates a new PostgreSQL role/user named `pentaho`
* `WITH PASSWORD 'SecurePassword123'` - Sets the password for this user

***

* `ALTER USER pentaho` - Modifies the existing `pentaho` user
* `WITH SUPERUSER` - Grants full superuser privileges (can do anything: create databases, create users, bypass all permissions, etc.)
* `# Demo only` - **Warning comment** indicating this is dangerous for production
  {% endhint %}

3. Test connection as `pentaho`:

```bash
sudo -u pentaho psql -d postgres -c "\conninfo"
```

{% endstep %}

{% step %}
**Optional: Allow remote connections**

{% hint style="info" %} <mark style="color:green;">For reference only as connecting to Postgresql via localhost</mark>

By default, PostgreSQL accepts only local connections. For localhost setups, you do not need this step.
{% endhint %}

1. Back up configs and edit `postgresql.conf` (PostgreSQL 17):

```bash
sudo cp /etc/postgresql/17/main/postgresql.conf{,.bak}
sudo nano /etc/postgresql/17/main/postgresql.conf
```

2. Set: `listen_addresses = '*'` (a specific interface/IP in production)

<figure><img src="/files/cQ1EvwJYt99JuzwazUqP" alt=""><figcaption><p>Set listener</p></figcaption></figure>

3. Save:

```
Ctrl + o
Enter
Ctrl + x
```

{% endstep %}

{% step %}
**Configure pg\_hba.conf**

{% hint style="info" %}
The **pg\_hba.conf** file (PostgreSQL Host-Based Authentication configuration) controls who can connect to your PostgreSQL database and how they authenticate. You configure it during installation to define security rules for database access.

In production you would modify the settings to only allow the required users access to the databases on Server IPs.
{% endhint %}

1. Configure PostgreSQL to use md5 password authentication `pg_hba.conf` .

```bash
sudo cp /etc/postgresql/17/main/pg_hba.conf{,.bak}
sudo nano /etc/postgresql/17/main/pg_hba.conf
```

2. Manually edit the file:

<figure><img src="/files/fb3fNbEzWNwogOfHDXkn" alt=""><figcaption></figcaption></figure>

<mark style="color:$primary;">Or</mark>

2. Run the following script.

```bash
# Use the exact path
PG_HBA_PATH=/etc/postgresql/17/main/pg_hba.conf

# Verify the file exists
ls -la "$PG_HBA_PATH"

# If that works, then run:
sudo cp "$PG_HBA_PATH" "${PG_HBA_PATH}.backup"
sudo sed -i 's/^local[[:space:]]\+all[[:space:]]\+postgres[[:space:]]\+peer$/local   all             postgres                                scram-sha-256/' "$PG_HBA_PATH"
sudo sed -i 's/^local[[:space:]]\+all[[:space:]]\+all[[:space:]]\+peer$/local   all             all                                     scram-sha-256/' "$PG_HBA_PATH"
sudo sed -i '0,/^host[[:space:]]\+all[[:space:]]\+all[[:space:]]\+127\.0\.0\.1\/32[[:space:]]\+scram-sha-256$/s//host    all             all             0.0.0.0\/0               md5/' "$PG_HBA_PATH"

# Verify
sudo cat "$PG_HBA_PATH" | grep -E "^(local|host)[[:space:]]+(all|replication)"
```

<figure><img src="/files/RcH06r3PBn3pdiJG0ufZ" alt=""><figcaption><p>SED: pg_hba.conf</p></figcaption></figure>

{% hint style="info" %}
Prefer `scram-sha-256` over `md5` in `pg_hba.conf` for stronger password hashing on modern PostgreSQL versions.

Ensure your JDBC driver supports SCRAM (e.g., recent PostgreSQL drivers). If compatibility issues arise, use `md5` as a fallback.
{% endhint %}

3. Restart the PostgreSQL service.

```bash
cd
systemctl restart postgresql
# Password: password
```

4. Allow firewall (if enabled) and restart:

```bash
sudo ufw allow 5432/tcp || true
sudo systemctl restart postgresql
sudo ss -ltnp | grep 5432
```

{% endstep %}

{% step %}
**Install pgAdmin 4 (desktop)**

Install the pgAdmin 4 desktop client using the official repository.

1. Add repo and key:

```bash
sudo install -d -m 0755 /etc/apt/keyrings
curl -fsS https://www.pgadmin.org/static/packages_pgadmin_org.pub | gpg --dearmor | sudo tee /etc/apt/keyrings/packages-pgadmin-org.gpg > /dev/null
source /etc/os-release
echo "deb [signed-by=/etc/apt/keyrings/packages-pgadmin-org.gpg] https://ftp.postgresql.org/pub/pgadmin/pgadmin4/apt/${UBUNTU_CODENAME} pgadmin4 main" | sudo tee /etc/apt/sources.list.d/pgadmin4.list
```

2. Install pgAdmin 4 (desktop):

```bash
sudo apt update -y && sudo apt upgrade -y
sudo apt install -y pgadmin4-desktop
```

3. Optional: verify repo entry

```bash
cat /etc/apt/sources.list.d/pgadmin4.list
```

4. Add a server connection in pgAdmin:

* Right‑click Servers → Register → Server
* Name: `Pentaho`
* Connection: host, port `5432`, user `pentaho`, password `SecurePassword123` (do not save in production)

<figure><img src="/files/uRCPhHzXSs1XqGJhaZc4" alt="Create server group dialog in pgAdmin"><figcaption><p>Create server group</p></figcaption></figure>

<figure><img src="/files/Nj1JY6CZiAIvM0NJSWH8" alt="pgAdmin server group list"><figcaption><p>Server group</p></figcaption></figure>

<figure><img src="/files/wXVhdH6uLVpjeRAaRbYh" alt="pgAdmin new server connection dialog"><figcaption><p>Connection details</p></figcaption></figure>

{% hint style="danger" %}
Do not save passwords in production. Prefer OS‑level keyrings or secure vaults.
{% endhint %}

<figure><img src="/files/t98tQiU6zh3KMGBSNq9X" alt="pgAdmin 4 main UI window"><figcaption><p>pgAdmin 4 UI</p></figcaption></figure>
{% endstep %}

{% step %}
**Validate the environment**

1. Quick checks to confirm everything is ready for the next step.

```bash
# Java
java -version
[ "$PENTAHO_JAVA_HOME" = "/usr/lib/jvm/java-21-openjdk-amd64" ] && echo OK || echo "Check PENTAHO_JAVA_HOME"

# PostgreSQL
sudo -u postgres psql -c "SELECT version();"
sudo systemctl status postgresql --no-pager | sed -n '1,5p'

```

<figure><img src="/files/7em8EiKR2hS1hVGgfmas" alt=""><figcaption><p>Validate PostgreSQL 17 service ..</p></figcaption></figure>
{% endstep %}
{% endstepper %}

***


# Install Pentaho Server

Installation of Pentaho Server components ..

{% hint style="info" %}

#### **Pentaho Server**

This section guides you through installing and starting the Pentaho Server on Ubuntu.

You will:

* Create installation directories
* Prepare the Pentaho Repository databases
* Configure JDBC/JNDI connections
* Start the Pentaho Server (and optionally set up systemd)
* Configure the License Manager
  {% endhint %}

{% hint style="warning" %}
Tested baseline: Ubuntu 24.04 LTS with Java 21 (OpenJDK) and PostgreSQL 17.

Make sure you have completed [Prepare Environment](/pentaho-11-installation-en/installation/archive-installation/prepare-environment) first.

For compatibility details, see [Components Reference](https://docs.pentaho.com/install/components-reference).
{% endhint %}

{% hint style="info" %}
**Prerequisites**

* Ubuntu 24.04 LTS server
* Java 21 installed and `PENTAHO_JAVA_HOME` set
* PostgreSQL 17 installed and running
* A non‑root `pentaho` user with sudo
* `unzip` package installed
* Archive ZIPs and JDBC drivers downloaded
  {% endhint %}

<figure><img src="/files/lIPH3nfqIrFj5qeQcx5F" alt="Pentaho Pro Suite overview image"><figcaption><p>Pentaho Pro Suite</p></figcaption></figure>

{% tabs %}
{% tab title="1. Pentaho Server" %}
{% hint style="info" %}

#### **Pentaho Server Directories**

The Pentaho Server is a web application running in a Apache Tomcat servlet container.
{% endhint %}

1. Create base directories under `/opt/pentaho`.

```bash
cd
sudo mkdir -p /opt/pentaho/{server,software}
```

```
/opt/pentaho
├── server      # Unpacked server runtime (tomcat, pentaho-solutions, scripts)
└── software    # Installers, ZIPs, drivers (staging area)
```

2. Create sub-directories in `/opt/pentaho/software`.

```bash
cd /opt/pentaho/software
sudo mkdir -p {db-drivers,docker,ee-client,ee-plugins,ops-mart,sdk,server,shims}
```

```
db-drivers   - JDBC drivers
docker       - Pentaho on-prem config & dockerfiles
ee-client    - Pentaho EE client plugins
ee-plugins   - Pentaho EE server plugins
ops-mart     - Operations Mart scripts
sdk          - Pentaho SDK kit
server       - Pentaho server
shims        - Hadoop shims collections
```

***

{% hint style="info" %}
**Unpack Pentaho Server Package (ZIP)**

Use `unzip` to extract the server ZIP into the runtime directory. This avoids requiring the full JDK (the JRE does not include the `jar` tool).

* `pentaho-server-ee-11.0.0.0-2xx.zip` - Pentaho Server (Archive - incl Tomcat 10)
  {% endhint %}

1. Ensure `unzip` is available and copy the server ZIPs into staging.

```bash
cd
sudo apt update -y && sudo apt install -y unzip
sudo cp ~/Downloads/'Archive Build (Suggested Installation Method)'/* /opt/pentaho/software/server
```

2. Extract the Pentaho Server ZIP into `$PENTAHO_BASE/server`.

```bash
cd
cd "$PENTAHO_BASE/server"

# Replace <version> with the exact file name you downloaded
sudo unzip /opt/pentaho/software/server/pentaho-server-ee-11.0.0.0-2xx.zip
```

<figure><img src="/files/1rLrfCVzT33Nwxij2LhJ" alt=""><figcaption><p>Unzip Pentaho Server</p></figcaption></figure>

3. Make all `.sh` files executable.

```bash
cd
cd "$PENTAHO_BASE/server"
sudo find . -iname "*.sh" -exec chmod +x {} \;
```

4. Set ownership and sensible permissions to run 'pentaho' as a non-root user.

```bash
cd
sudo chown -R pentaho:pentaho /opt/pentaho
sudo find /opt/pentaho -type d -exec chmod 755 {} \;
sudo find /opt/pentaho -type f -exec chmod 644 {} \;
sudo find /opt/pentaho -name "*.sh" -exec chmod 755 {} \;
```

{% hint style="info" %}
755 means you can do anything with the file or directory, and other users can read and execute it but not alter it. Suitable for programs and directories you want to make publicly available.

644 means you can read and write the file or directory and other users can only read it.
{% endhint %}

5. Verify the server directory structure.

{% hint style="info" %}
/opt/pentaho/

```
  server/
    pentaho-server/
      pentaho-solutions/
        system/
```

Server plugins are installed into the `pentaho-solutions/system` folder.
{% endhint %}
{% endtab %}

{% tab title="2. Pentaho Repository" %}
{% hint style="info" %}

#### **Pentaho Repository components**

The Pentaho Repository (on PostgreSQL by default) consists of:

* Jackrabbit: solution repository, security and content metadata
* Quartz: scheduler data
* Hibernate: audit logging
* Pentaho Operations Mart: usage and performance reporting
  {% endhint %}

{% stepper %}
{% step %}
**Review default passwords in SQL scripts**

1. Inspect the PostgreSQL scripts shipped with the server.

```bash
cd
cd "$PENTAHO_SERVER/data/postgresql"
ls -1
```

You should see files similar to:

```
alter_script_postgresql_BISERVER-13674.sql
create_jcr_postgresql.sql
create_quartz_postgresql.sql
create_repository_postgresql.sql
migrate_old_quartz_data_postgresql.sql
pentaho_logging_postgresql.sql
pentaho_mart_drop_postgresql.sql
pentaho_mart_postgresql.sql
pentaho_mart_upgrade_audit_postgresql.sql
pentaho_mart_upgrade_postgresql.sql

```

2. Open a script to review default users/passwords (change for production).

```bash
sed -n '1,120p' create_jcr_postgresql.sql
```

{% hint style="danger" %}
Production guidance: use unique, strong passwords, rotate them regularly, store them in a secure vault, and never commit secrets to source control.

Do not keep workshop password defaults in production.
{% endhint %}
{% endstep %}

{% step %}
**Run SQL scripts to create Repository databases**

1. Confirm PostgreSQL is running and locate the scripts.

```bash
sudo systemctl status postgresql --no-pager
cd
cd "$PENTAHO_SERVER/data/postgresql"
ls -l
```

2. Connect as a `pentaho` superuser.

Password: `SecurePassword123`

```bash
sudo -u pentaho psql -d postgres
```

3. Execute the commands step-by-step - not as a single script block. Provide passwords if prompted - see below.

```plsql
\i create_jcr_postgresql.sql
\i create_quartz_postgresql.sql
\q quit after Quartz and log back in ..
ensure your in the "$PENTAHO_SERVER/data/postgresql" directory
sudo -u pentaho psql -d postgres
\i create_repository_postgresql.sql
\i pentaho_mart_postgresql.sql
\q quit after hibernate and log back in ..
ensure your in the "$PENTAHO_SERVER/data/postgresql" directory
sudo -u pentaho psql -d postgres
\i pentaho_logging_postgresql.sql
\q
```

| User          | Password          |
| ------------- | ----------------- |
| postgres      | SecurePassword123 |
| pentaho       | SecurePassword123 |
| jcr\_user     | password          |
| pentaho\_user | password          |
| hibuser       | password          |

4. Quick validation (CLI) - list created databases and connect - hit q to scroll through list.

```sql
\l+ jackrabbit
\l+ quartz
\l+ hibuser
\l+ opsmart
\c jackrabbit
\dt
\c quartz
\dt
```

{% hint style="warning" %}
Tables for Hibernate and Jackrabbit may be created later by the Pentaho Server on first start. Seeing empty schemas at this stage can be expected.
{% endhint %}

5. Optional: verify in pgAdmin (GUI).

<figure><img src="/files/B9XKROELAPp4VFMpkTS6" alt=""><figcaption><p>Pentaho databases</p></figcaption></figure>
{% endstep %}

{% step %}
**Configure Pentaho to use PostgreSQL**

{% hint style="info" %}
PostgreSQL is the default. If you kept the default passwords and port (`5432`), only verify the settings below; otherwise, adjust host/port/user/password to match your environment.
{% endhint %}

***

{% hint style="info" %}
Quartz - set PostgreSQL delegate and JNDI data source.
{% endhint %}

1. Open the Quartz configuration.

```bash
cd "$PENTAHO_SERVER/pentaho-solutions/system/scheduler-plugin/quartz"
sudo nano -c quartz.properties
```

2. Check these values (line numbers may differ):

{% code title="quartz.properties — required entries" %}

```
org.quartz.jobStore.driverDelegateClass=org.quartz.impl.jdbcjobstore.PostgreSQLDelegate
org.quartz.dataSource.myDS.jndiURL=Quartz
```

{% endcode %}

***

{% hint style="info" %}
Hibernate — point to the PostgreSQL configuration file.
{% endhint %}

1. Open Hibernate settings.

```bash
cd "$PENTAHO_SERVER/pentaho-solutions/system/hibernate"
sudo nano -c hibernate-settings.xml
```

2. Confirm the config file reference:

{% code title="hibernate-settings.xml" %}

```
<config-file>system/hibernate/postgresql.hibernate.cfg.xml</config-file>
```

{% endcode %}

3. Optionally review `postgresql.hibernate.cfg.xml` for datasource name and dialect.

```bash
sudo nano -c postgresql.hibernate.cfg.xml
```

Ensure:

{% code title="postgresql.hibernate.cfg.xml — key properties" %}

```
<property name="connection.driver_class">org.postgresql.Driver</property>
<property name="dialect">org.hibernate.dialect.PostgreSQLDialect</property>
<property name="hibernate.connection.datasource">java:comp/env/jdbc/Hibernate</property>
```

{% endcode %}

***

{% hint style="info" %}
Jackrabbit — verify PostgreSQL storage in `repository.xml`.
{% endhint %}

1. Open Jackrabbit configuration.

```bash
cd "$PENTAHO_SERVER/pentaho-solutions/system/jackrabbit"
sudo nano -c repository.xml
```

2. Check that PostgreSQL sections are active (others commented out), for example:

{% hint style="info" %}

* Filesystem schema: `postgresql`
* Datastore `databaseType="postgresql"`
* PersistenceManager schema: `postgresql`
* Database Journal: `postgresql`
  {% endhint %}

{% hint style="warning" %}
Cross‑check JNDI names and ports across Quartz, Hibernate, Jackrabbit, and Tomcat `context.xml` to ensure consistency (same host, port 5432 unless changed, and matching JNDI resource names).

Expected JNDI names:

* Quartz: `Quartz`
* Hibernate: `java:comp/env/jdbc/Hibernate`
* Jackrabbit: as referenced in `repository.xml`
  {% endhint %}
  {% endstep %}
  {% endstepper %}
  {% endtab %}

{% tab title="3. Tomcat" %}
{% hint style="info" %}

#### **Tomcat**

After configuring the Repository, configure the web application server (Tomcat 10) to connect to the Repository using JDBC/JNDI.
{% endhint %}

{% tabs %}
{% tab title="1. Database Drivers" %}
{% hint style="warning" %}

#### **JDBC Drivers**

To connect to databases (including the Repository), install the appropriate JDBC drivers. Due to licensing restrictions, some drivers must be downloaded manually.
{% endhint %}

{% embed url="<https://docs.pentaho.com/install/jdbc-drivers-reference>" %}

{% embed url="<https://jdbc.postgresql.org/download/postgresql-42.7.8.jar>" %}

1. Verify the PostgreSQL JDBC driver is present in Tomcat lib (required for the Repository):

```bash
ls -1 "$TOMCAT_HOME/lib" | grep -i postgresql || echo "PostgreSQL driver not found"
```

{% hint style="info" %}
If not found, download the PostgreSQL JDBC driver (e.g., `postgresql-42.7.8.jar`) and distribute it using the helper script:
{% endhint %}

```bash
sudo cp ~/Downloads/'Database Drivers'/postgresql-*.jar /opt/pentaho/server/jdbc-distribution
cd /opt/pentaho/server/jdbc-distribution
sudo ./distribute-files.sh "$TOMCAT_HOME/lib"
```

2. Copy any additional JDBC drivers to the staging folder and distribute.

```bash
sudo cp ~/Downloads/'Database Drivers'/* /opt/pentaho/software/db-drivers
cd /opt/pentaho/software/db-drivers
sudo cp mysql-connector-j-9.0.0.jar /opt/pentaho/server/jdbc-distribution
```

3. Distribute the drivers to Tomcat.

```bash
cd /opt/pentaho/server/jdbc-distribution
sudo ./distribute-files.sh "$TOMCAT_HOME/lib"
```

4. Verify the JARs are present in Tomcat lib.

```bash
ls -1 "$TOMCAT_HOME/lib" | grep -Ei 'mysql|postgresql' || echo "Not found"
```

{% hint style="danger" %}
You must restart the Pentaho Server (and client tools, if running) to load new JDBC drivers. A full system reboot is not required.
{% endhint %}
{% endtab %}

{% tab title="2. context.xml" %}
{% hint style="info" %}

#### **context.xml**

Database connection information for JNDI resources used by Pentaho is stored in `context.xml`.
{% endhint %}

1. Open the file and review JNDI resources and credentials.

```bash
cd "$TOMCAT_HOME/webapps/pentaho/META-INF"
sudo nano -c context.xml
```

{% hint style="warning" %}
In production, verify username, password, driver class, host/IP, and port match your environment. Ensure JNDI names align with Quartz, Hibernate, and Jackrabbit configurations.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="4. Start Server" %}
{% hint style="info" %}

#### **Start Pentaho Server**

Run the server as the `pentaho` user to avoid permission issues.
{% endhint %}

1. Start the server.

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

2. Tail the Tomcat log (robust pattern).

```bash
tail -f "$TOMCAT_HOME/logs/catalina."*.log
```

Expected messages include:

```
... Starting ProtocolHandler ["http-nio-8080"]
... Server startup in [xxxxx] milliseconds
```

***

{% hint style="info" %}
**Pentaho User Console (PUC)**

The Pentaho User Console (PUC) is the web UI for creating and viewing content.
{% endhint %}

{% embed url="<http://localhost:8080/pentaho>" %}
Link to Pentaho Server
{% endembed %}

Default credentials (change immediately after first login):

```
Username: admin
Password: password
```

<figure><img src="/files/pEiMXzquH8k27qVFdCmK" alt=""><figcaption><p>Pentaho User Console</p></figcaption></figure>

3. You also have the option to switch to the new login screen.

{% embed url="<http://localhost:8080/pentaho/content/login/web/index.html>" %}

<figure><img src="/files/SYe00G5Czh0wdt8HZvyX" alt=""><figcaption><p>NEW - Pentaho User Console</p></figcaption></figure>

{% hint style="info" %}
If you have already entered your licensing details, then you will be redirected to the new User Console.
{% endhint %}

***

{% hint style="success" %}
**Quick validations**
{% endhint %}

* Check HTTP is responding:

```bash
curl -I http://localhost:8080/pentaho/ | head -n 1
```

* Optional (remote access): open firewall and test from a client machine (adjust to your network policy):

```bash
sudo ufw allow 8080/tcp || true
```

***

<details>

<summary>Validate Repository after first run (click to expand)</summary>

```sql
-- Quartz example
\c quartz
SELECT COUNT(*) FROM qrtz_scheduler_state;

-- Hibernate example (expect many tables after first start)
\c hibuser
\dt
```

</details>

***

{% hint style="info" %}
**Systemd (optional)**

Create a systemd service to manage Pentaho at boot and on failure.
{% endhint %}

1. Save the unit file as `/etc/systemd/system/pentaho-server.service`.

{% code title="/etc/systemd/system/pentaho-server.service" %}

```ini
[Unit]
Description=Pentaho Server
After=network-online.target
Wants=network-online.target

[Service]
Type=forking
User=pentaho
Group=pentaho
EnvironmentFile=-/etc/environment
# Optional fallback if not using PENTAHO_JAVA_HOME in /etc/environment
# Environment="JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64"
WorkingDirectory=/opt/pentaho/server/pentaho-server
ExecStart=/opt/pentaho/server/pentaho-server/start-pentaho.sh
ExecStop=/opt/pentaho/server/pentaho-server/stop-pentaho.sh
TimeoutSec=500
Restart=on-failure
RestartSec=5
SuccessExitStatus=5 6
LimitNOFILE=65535

[Install]
WantedBy=multi-user.target
```

{% endcode %}

2. Reload and start.

```bash
sudo systemctl daemon-reload
sudo systemctl start pentaho-server
sudo systemctl enable pentaho-server
```

Manage the service:

```bash
sudo systemctl status pentaho-server
sudo systemctl restart pentaho-server
sudo systemctl stop pentaho-server
```

***

<details>

<summary>Troubleshooting (click to expand)</summary>

* HTTP 404 on `/pentaho` after startup: confirm `"$TOMCAT_HOME/webapps/pentaho"` exists, check `catalina.*.log` for deployment errors, and verify file permissions under `$PENTAHO_SERVER`.
* Port 8080 already in use: change Tomcat port in `server.xml` or stop the conflicting service.
* JDBC driver not found: verify the driver JARs (e.g., `postgresql-*.jar`, `mysql-*.jar`) exist in `"$TOMCAT_HOME/lib"`.
* Authentication to PostgreSQL fails: prefer `scram-sha-256`; review `pg_hba.conf`, restart PostgreSQL, and test `psql -h 127.0.0.1` with the target user.
* Jackrabbit schema issues: re-check `repository.xml` sections are set to `postgresql`.
* License activation errors: check Tomcat logs in `tomcat/logs/` for `license`/`elm` messages.

</details>

***

{% endtab %}

{% tab title="5. License Manager" %}
{% hint style="info" %}

#### **Licensing Manager**

Pentaho Pro Suite 11.x uses a License Manager (cloud or local) to manage PDI & BA entitlements and verify EE plugins.
{% endhint %}

<figure><img src="/files/82g7LapAxmEtvoJQFN0s" alt=""><figcaption><p>Licenses</p></figcaption></figure>

{% tabs %}
{% tab title="License Manager" %}
{% hint style="info" %}

#### **Trial license**

A 30-day trial license is included if you have downloaded from: [Pentaho 30-day Trial](https://pentaho.com/download/)

If you have dowwnloaded the GA binaries from: [Pentaho Customer Portal](https://support.pentaho.com/hc/en-us), then you will require an Activation ID or your LIcensing URL.

If you have installed in an air-gapped envirnoment, you will need to request an offline license.
{% endhint %}

1. Launch Pentaho Server > Administration > Licenses to open the Add License dialog.
2. Click the + sign.

<figure><img src="/files/l1mDtUtxQBtxM6SYiixI" alt=""><figcaption><p>Add license</p></figcaption></figure>

5. Enter Activation code or your licensing URL:

<figure><img src="/files/Jfa9LcSdpDL1VklTZBzU" alt="Add License dialog"><figcaption><p>License Manager</p></figcaption></figure>

{% hint style="warning" %}
**Enterprise licenses**

If upgrading from 9.x or earlier, install the new product version before activating licenses. Do not start the server before upgrading the licenses.
{% endhint %}
{% endtab %}

{% tab title="Set ENV License Path" %}
{% hint style="info" %}
**Set license path environment variable**

Create a `PENTAHO_LICENSE_INFORMATION_PATH` environment variable so the Pentaho Server consistently finds your license file.
{% endhint %}

1. Ensure the target directory exists and is secured.

```bash
sudo -u pentaho mkdir -p /home/pentaho/.pentaho
sudo chown -R pentaho:pentaho /home/pentaho/.pentaho
chmod 700 /home/pentaho/.pentaho
```

2. Edit `/etc/environment` and add the line below (no `export`).

```bash
sudo nano /etc/environment
```

Append (or update) the following:

```
PENTAHO_LICENSE_INFORMATION_PATH=/home/pentaho/.pentaho/.elmLicInfo.plt
```

3. Log out/log in or reload the environment and verify.

```bash
source /etc/environment
env | grep PENTAHO_LICENSE_INFORMATION_PATH
```

The `PENTAHO_LICENSE_INFORMATION_PATH` variable is now set.
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Server Plugins

Installation of Pentaho User Console  Plugins ..

{% hint style="info" %}

#### Plugin Manager

To make deployment of the Pentaho Server plugins easier there's a modern PUC with a built-in Plugin Manager.

From the Plugin Manager UI you can now manage your plugin lifecycle:

* Update Available - check for updates
* Installed - list installed plugins
* Not Installed - list plugins not installed

for both Server and Client side EE plugins.
{% endhint %}

***

1. Log in:

{% embed url="<https://localhost:8080/pentaho/content/login/web/index.html>" %}

2. Select Plugin Manager.

<figure><img src="/files/OWVZTqeePwHVZJajV9Jy" alt=""><figcaption><p>Plugin Manager</p></figcaption></figure>

Or

1. Log in & Switch to Modern Design.

<figure><img src="/files/X0hxfA9PE6Ba2YpT7ch0" alt=""><figcaption><p>Switch to Modern Design</p></figcaption></figure>

{% tabs %}
{% tab title="1. Plugin Manager" %}
{% hint style="info" %}

#### **NEW - Pentaho User Console**

The left panel displays the main navigation options including Home (currently active), Browse Files, Plugin Manager, Scheduler, Data Connections, Settings, Semantic Model Editor, and Pipeline Designer.

The Home page features a Quick Access section with four tiles: Data Sources for managing project data sources, Browse Files for exploring files to use with Pentaho, Semantic Model Editor for viewing or creating semantic models, and Pipeline Designer for creating transformations and jobs in the new web-based editor.

At the bottom, the Recently Opened section displays two .ktr files—tr\_write\_output and tr\_hello\_world (marked as favorite) - both last modified on December 12, 2025, and owned by the admin user.
{% endhint %}

<figure><img src="/files/cF8k14HjNmw65M14UsaI" alt=""><figcaption><p>NEW - Pentaho User Console</p></figcaption></figure>

{% tabs %}
{% tab title="Analytic Plugins" %}
{% hint style="info" %}

#### Analytic Plugins

Plugin Manager - The recommended method for managing the lifecycle of your plugins.
{% endhint %}

{% tabs %}
{% tab title="1. Analyzer" %}
{% hint style="info" %}

#### **Analyzer**

Pentaho Analyzer is a web-based business intelligence tool that's part of the Pentaho Business Analytics platform. It provides an interactive, drag-and-drop interface for analyzing data and creating visualizations without requiring SQL or technical coding knowledge.

The tool allows users to explore data through OLAP (Online Analytical Processing) cubes, enabling multidimensional analysis with features like drill-down, slice-and-dice, and pivot operations. Users can quickly create charts, graphs, and reports by dragging dimensions and measures onto a canvas, making it accessible for business users who need to perform ad-hoc analysis.

Pentaho Analyzer supports various visualization types including bar charts, line graphs, pie charts, and heat grids, and integrates with the broader Pentaho platform for sharing reports and embedding analytics into dashboards. It's particularly useful for organizations that want to empower business users to independently explore and visualize their data warehouse or mart information.
{% endhint %}

1. Select: Analyzer

<figure><img src="/files/68XSsFmRWcPuxQeWXKxe" alt=""><figcaption><p>Analyzer</p></figcaption></figure>

2. From the drop-down box, select : Version

<figure><img src="/files/nCpJo79DKeMwCEkRv4HU" alt=""><figcaption><p>Analyzer Plugin</p></figcaption></figure>

3. Click Install.
4. Optional: Verify Analyzer is installed.

```bash
[ -d analyzer ] && echo OK || echo "Analyzer directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}
{% endtab %}

{% tab title="2. Interactive Reporting" %}
{% hint style="info" %}

#### **Interactive Reporting**

Pentaho Interactive Reporting (PIR) is a web-based ad-hoc reporting tool within the Pentaho Business Analytics platform that enables users to create and customize reports through an intuitive interface without requiring technical expertise.

The tool provides a WYSIWYG (What You See Is What You Get) drag-and-drop environment where users can build reports by selecting data sources, adding fields, applying filters, and formatting output. Unlike traditional report design tools that require developer skills, PIR is designed for business users who need to quickly generate operational reports and answer specific business questions.

Key capabilities include the ability to create tabular reports with grouping, sorting, filtering, and calculated fields. Users can add charts, apply conditional formatting, and create prompts for parameterized reports. The tool supports various output formats including HTML, PDF, Excel, and CSV, making it easy to distribute reports across the organization.

PIR connects to relational databases and Pentaho data sources, allowing users to work with live data. It's particularly valuable for organizations that want to democratize reporting capabilities and reduce the bottleneck of relying solely on IT or developers to create standard operational reports.
{% endhint %}

1. Select: Interactive Reporting.

<figure><img src="/files/S1E361z8w8aBY4G2Kuwh" alt=""><figcaption><p>Interactive Reporting</p></figcaption></figure>

<figure><img src="/files/LBWsJwaNIfs2Dgzpmyvm" alt=""><figcaption><p>Interactive Reporting</p></figcaption></figure>

2. From the drop-down box, select : Version

<figure><img src="/files/ah5IP08l1MABNrCTUyb0" alt=""><figcaption><p>Interactive Reporting Plugin</p></figcaption></figure>

2. Click Install.
3. Optional: Verify Interactive Reporting is installed.

```bash
[ -d pentaho-interactive ] && echo OK || echo "Interactive Reporting directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}
{% endtab %}

{% tab title="3. Dashboard Designer" %}
{% hint style="info" %}

#### **Dashboard Designer**

Pentaho Dashboard Designer (also known as CDE - Community Dashboard Editor in the community edition) is a web-based tool for creating interactive, customizable dashboards within the Pentaho Business Analytics platform.

The tool provides a comprehensive environment for building dashboards that combine multiple visualizations, reports, and interactive components into a single unified interface. Users can create dashboards by assembling various elements including charts, tables, filters, selectors, and other widgets that work together to provide a complete analytical experience.

Dashboard Designer uses a three-panel approach: Layout (for defining the dashboard structure using HTML/CSS), Components (for adding data-driven elements like charts and tables), and Data Sources (for connecting to queries and data). This architecture allows for significant flexibility and customization, though it does require some technical knowledge, particularly for advanced layouts and styling.

Key features include interactivity through parameter passing between components, allowing filters and selectors to dynamically update multiple visualizations simultaneously. The tool supports various charting libraries (CCC - Community Chart Components being the most common), and can integrate content from other Pentaho tools like Analyzer and Interactive Reporting.

Pentaho Dashboard Designer is particularly useful for creating executive dashboards, operational monitoring screens, and analytical applications where users need to see multiple related metrics and visualizations in context, with the ability to drill down and filter data interactively.
{% endhint %}

1. Select: Dashboard Designer.

<figure><img src="/files/2GhcylAfpNsm00LrnDpY" alt=""><figcaption><p>Dashboard Designer</p></figcaption></figure>

2. From the drop-down box, select : Version

<figure><img src="/files/KPjCc3UpZt1CE29gv1Ph" alt=""><figcaption><p>Dashboard Designer Plugin</p></figcaption></figure>

3. Click Install.
4. Optional: Verify Dashboard Designer is installed.

```bash
[ -d dashboards ] && echo OK || echo "Dashboard Designer directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="NEW Plugins" %}
{% hint style="info" %}

#### Pentaho 11 Server Plugins

{% endhint %}

Browse the various plugins:

{% tabs %}
{% tab title="1. Pipeline Designer" %}
{% hint style="info" %}

#### **Pipeline Designer**

The **Pentaho Pipeline Designer** is a visual tool that lets you build data transformation workflows through a drag-and-drop interface. The left panel contains a searchable library of pre-built components (like CSV inputs, MongoDB operations, data generators), the center canvas is where you visually connect these components into a flow diagram to create your pipeline, and the bottom logging console shows real-time execution details and performance metrics when you run the transformation.

It's essentially a no-code environment for designing data pipelines, and it's one of the plugin components that gets deployed in your Pentaho Server setup.
{% endhint %}

1. Select: Pipeline Designer.

<figure><img src="/files/bFWvfpiZCfyoFUBmhT8d" alt=""><figcaption><p>Pipeline Designer</p></figcaption></figure>

2. From the drop-down box, select : Version

<figure><img src="/files/QXDocYuofSSY2T6pNepR" alt=""><figcaption><p>Pipeline Designer Plugin</p></figcaption></figure>

3. Click Install.
4. Optional: Verify Pipeline Designer is installed.

```bash
[ -d pentaho-webttle ] && echo OK || echo "Pipeline Designer directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

6. Log in:

{% embed url="<https://localhost:8080/pentaho/content/login/web/index.html>" %}

7. Select: Pipeline Designer.

<figure><img src="/files/jUEVfgZKkZVYdSjm0roV" alt=""><figcaption><p>Piopeline Designer</p></figcaption></figure>

8. Create a test Transformation.

<figure><img src="/files/QVfHpQipUyNoEt9UBB36" alt=""><figcaption><p>Transformation - Hello World</p></figcaption></figure>
{% endtab %}

{% tab title="2. Semantic Model Editor" %}
{% hint style="info" %}

#### **Semantic Model Editor**

The Semantic Model Editor (SME) helps you create data models that define how your data should be organized and analyzed for business intelligence. It defines a semantic layer between your raw data and your reports, ensuring everyone in your organization uses consistent business logic and definitions.
{% endhint %}

1. Select: Semantic Model Editor.

<figure><img src="/files/oXPqLbN0k3TsNIcLPfRu" alt=""><figcaption><p>Semantic Model Editor</p></figcaption></figure>

2. From the drop-down box, select : Version

<figure><img src="/files/ysACWzMs9H6Ot6X0IHQt" alt=""><figcaption></figcaption></figure>

3. Click Install.
4. Optional: Verify semantic model editor is installed.

```bash
[ -d semantic-model-editor ] && echo OK || echo "semantic-model-editor directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

6. Log in:

{% embed url="<https://localhost:8080/pentaho/content/login/web/index.html>" %}

7. Select: Model Editor.

<figure><img src="/files/AGT8Lnv80u83lpByLJUz" alt=""><figcaption><p>Semantic Model Editor</p></figcaption></figure>

<figure><img src="/files/H65BsZSmspI4haqQPDWq" alt=""><figcaption><p>SteelWheels Model</p></figcaption></figure>
{% endtab %}

{% tab title="3. Scheduler" %}
{% hint style="info" %}

#### **Scheduler**

The **Pentaho Scheduler** provides centralized management and automation of scheduled tasks across your Pentaho system. It displays a table of all scheduled jobs showing their source files, execution frequency (like "Every day at 3:15 PM"), current status (active/paused), last run times, and output locations.

You can view all schedules or filter by active/paused status, manually execute jobs on-demand, set blockout times when jobs shouldn't run, and pause/resume the entire scheduler - essentially giving you complete visibility and control over automated data transformation workflows and their execution timing.
{% endhint %}

1. Stop Pentaho Server.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

2. Confirm the plugin archive exists.

```bash
ls -1 "$PENTAHO_BASE/software/ee-plugins" | grep -i pas-scheduler || echo "Pipeline Scheduler plugin ZIP not found"
```

3. Extract `pas-scheduler-*.zip` into the `system` folder.

```bash
cd
cd "$PENTAHO_SERVER/pentaho-solutions/system"
unzip "$PENTAHO_BASE/software/ee-plugins/pas-scheduler-11.0.0.0-204.zip"
```

4. Verify directory structure.

```bash
[ -d semantic ] && echo OK || echo "pas-scheduler directory missing"
```

5. Start Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

6. Log in:

{% embed url="<https://localhost:8080/pentaho/content/login/web/index.html>" %}

7. Select: Scheduler

<figure><img src="/files/LsHA6cBDMxvpOVmnJO78" alt=""><figcaption><p>Scheduler</p></figcaption></figure>
{% endtab %}

{% tab title="4. Carte" %}
{% hint style="info" %}

#### **Pipeline Carte Server**

The **Pentaho Carte Server** is the execution engine that runs and monitors Pentaho transformations and jobs. It provides a web-based status dashboard showing all running and completed transformations/jobs with their execution details (status, timestamps, unique IDs), detailed step-by-step performance metrics (rows read/written, processing speed, errors), and configuration settings for log management and object lifecycle. The server tracks real-time execution stats for each transformation step, displays visual canvas previews of the pipeline flow, and maintains comprehensive execution logs - essentially serving as both the runtime engine and monitoring console for your workflows.
{% endhint %}

1. Verify directory structure.

```bash
[ -d webttle-carte-api-plugin ] && echo OK || echo " directory missing"
```

5. Restart Pentaho Server and verify in the UI.

```bash
cd
cd "$PENTAHO_SERVER"
./stop-pentaho.sh
```

```bash
cd
cd "$PENTAHO_SERVER"
./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}

6. RUN your test Transformation - see Pipeline Designer.

<figure><img src="/files/iLJNTwnki5bIUQsffZ6M" alt=""><figcaption><p>Carte Status</p></figcaption></figure>

<figure><img src="/files/mEV40ITKjjzxYB0N37uy" alt=""><figcaption><p>Carte details</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Data Connections" %}
{% hint style="info" %}

#### Data Connections

**Data Connections** manages data sources for the system.

The interface features a search bar for filtering data sources by name and an "Add connection" button in the upper right corner for configuring new data sources.

Each data source entry includes a checkbox for selection, displays the source name with an icon indicating the database type, and provides an "Open" button for accessing the data source along with additional configuration options accessible via a menu.
{% endhint %}

1. To view the connection details, Click: Open

<figure><img src="/files/CWh7GzSKkGUOfqr96CaD" alt=""><figcaption><p>SampleData connection</p></figcaption></figure>

2. The connection details panel enables you to 'tune' the connection.

<figure><img src="/files/8FaEw1GCZRJljuY7FlNy" alt=""><figcaption><p>SampleData connection details</p></figcaption></figure>

3. Create new will enable you to create a new database connection - in this example, Hypersonic.

<figure><img src="/files/UrQSs0vhw0sI3XkE2rmt" alt=""><figcaption><p>Create new database connection</p></figcaption></figure>

***

**Add Connection**

1. Click: Add Connection - in the main panel

<figure><img src="/files/XszOtU8EiMXxeKBDcxdd" alt=""><figcaption><p>Add Connection</p></figcaption></figure>

2. Click: Connect - to configure the connection.

<figure><img src="/files/WNGYeCTQZMgD0O5yjxti" alt=""><figcaption><p>Configure connection</p></figcaption></figure>

{% hint style="danger" %}
Remember to copy over the supported JDBC driver to:

/opt/pentaho/server/pentaho-server/tomcat/lib directory & restart the Pentaho server.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Plugin Matrix" %}

<table><thead><tr><th width="217">Plugin</th><th>Description</th></tr></thead><tbody><tr><td>Analyzer</td><td></td></tr><tr><td>Interactive Reporting</td><td></td></tr><tr><td>Dashboard</td><td></td></tr><tr><td>Pipeline Designer</td><td></td></tr><tr><td>Semantic Model Editor</td><td></td></tr><tr><td>Pipeline Designer</td><td></td></tr><tr><td>Pipeline Designer Carte</td><td></td></tr><tr><td>Elastic MapReduce</td><td></td></tr></tbody></table>

x
{% endtab %}
{% endtabs %}


# Install Client Tools

Installation of Clients Tools ..

{% hint style="info" %}

#### **Pentaho Client Tools**

There are two methods for installing the Business Analytics (BA) design tools. You can use either of these methods:

* Pentaho Business Analytics Evaluation Wizard — Windows Desktop
* Install each tool manually — Linux / Windows Desktop

The Evaluation Wizard is the easiest way to install design tools, utilities, or plugins on client workstations. Manual installation lets you place design tool files wherever needed. Choose a deployment that matches your DevOps and desktop constraints.
{% endhint %}

{% hint style="warning" %}
Baseline: Ubuntu 24.04 LTS with Java 21 (OpenJDK). For Windows, see: [Windows Installation](/pentaho-11-installation-en/installation/evaluation-installation).
{% endhint %}

<figure><img src="/files/lIPH3nfqIrFj5qeQcx5F" alt="Pentaho Pro Suite - client tools overview"><figcaption><p>Pentaho Pro Suite</p></figcaption></figure>

The following steps install the client tools on a Linux Desktop. ZIP filenames shown are examples; use your actual versions.

{% hint style="info" %}
**Unpack Pentaho Client Package (ZIP)**

Use `unzip` to extract the server ZIP into the runtime directory. This avoids requiring the full JDK (the JRE does not include the `jar` tool).

* `pdi-ee-11.0.0.0-237.zip` - Pentaho Data Integration

* `pad-ee-11.0.0.0-237.zip` - Pentaho Aggregation Designer

* `psw-ee-11.0.0.0-237.zip` - Pentaho Schema Workbench

* `pme-ee-11.0.0.0-237.zip` - Pentaho Metadata Editor

* `prd-ee-11.0.0.0-237.zip` - Pentaho Report Designer
  {% endhint %}

* Ensure `unzip` is installed:

```bash
sudo apt update -y && sudo apt install -y unzip
```

1. Create base Pentaho directory: \~`/Pentaho/design-tools`.

```bash
cd
mkdir -p ~/Pentaho/design-tools
```

<pre><code>~/Pentaho/design-tools
├── design-tools             
    └── data-integration  
    └── metadata-editor   
<strong>    └── schema-workbench      
</strong></code></pre>

{% tabs %}
{% tab title="1. Data Integration" %}
{% hint style="info" %}

#### **Pentaho Data Integration (PDI)**

Pentaho Data Integration (PDI) provides ETL capabilities for capturing, cleansing, and transforming data.
{% endhint %}

1. Locate `pdi-ee-11.0.0.0-237.zip`.

```bash
ls -1 ~/Downloads/'Client Tools'/'PDI (Spoon)'
```

3. Extract `pdi-ee-11.0.0.0-237.zip`

```bash
cd
cd ~/Pentaho/design-tools

# Replace <version> with the exact file name you downloaded.
unzip ~/Downloads/'Client Tools'/'PDI (Spoon)'/pdi-ee-11.0.0.0-2xx.zip
# You may need to adjust the path.
```

4. Make `.sh` files executable.

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec chmod +x {} \;
```

5. Verify structure:

{% hint style="info" %}
\~/Pentaho/design-tools/

* data-integration
* jdbc-distribution
* license-installer
  {% endhint %}

***

{% hint style="warning" %}
**Ubuntu 24.04 desktop UI notes**

Some legacy UI components in Spoon may require GTK/WebKit libraries that vary by desktop flavor. On Ubuntu 24.04, if Spoon reports missing GTK/WebKit modules, follow the instructions: [Missing GTK/WebKit modules](#missing-gtk-webkit-modules)
{% endhint %}

6. Start PDI (Spoon):

```bash
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

<figure><img src="/files/ypIzAR3guPYcSqkYxoPi" alt=""><figcaption><p>Spoon UI</p></figcaption></figure>
{% endtab %}

{% tab title="2. Metadata Editor" %}
{% hint style="info" %}

#### **Pentaho Metadata Editor (PME)**

Pentaho Metadata Editor creates and manages business-friendly semantic models.
{% endhint %}

1. Locate `pme-ee-11.0.0.0-237.zip`.

```bash
ls -1 ~/Downloads/'Client Tools'/'Metadata Editor'
```

2. Extract `pme-ee-11.0.0.0-237.zip`.

```bash
cd
cd ~/Pentaho/design-tools

# Replace <version> with the exact file name you downloaded.
unzip ~/Downloads/'Client Tools'/'Metadata Editor'/pme-ee-11.0.0.0-2xx.zip
# You may need to adjust the path.
```

3. Make `.sh` files executable.

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec chmod +x {} \;
```

4. Verify structure:

{% hint style="info" %}
\~/Pentaho/design-tools/

* data-integration
* jdbc-distribution
* license-installer
* metadata-editor
  {% endhint %}

5. Start PME:

```bash
cd
cd ~/Pentaho/design-tools/metadata-editor
./metadata-editor.sh
```

<figure><img src="/files/fygXWKMWQpuzV8spnreY" alt="Pentaho Metadata Editor UI"><figcaption><p>Pentaho Metadata Editor</p></figcaption></figure>
{% endtab %}

{% tab title="3. Schema Workbench" %}
{% hint style="info" %}

#### **Schema Workbench (PSW)**

Schema Workbench is used to create and test Mondrian OLAP cube schemas.
{% endhint %}

1. Locate `psw-ee-11.0.0.0-237.zip`.

```bash
ls -1 ~/Downloads/'Client Tools'/'Schema Workbench'
```

2. Extract `psw-ee-11.0.0.0-237.zip`.

```bash
cd
cd ~/Pentaho/design-tools

# Replace <version> with the exact file name you downloaded.
unzip ~/Downloads/'Client Tools'/'Schema Workbench'/psw-ee-10.2.0.0-2xx.zip
# You may need to adjust the path.
```

2. Make `.sh` files executable.

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec chmod +x {} \;
```

3. Verify structure:

{% hint style="info" %}
\~/Pentaho/design-tools/

* data-integration
* jdbc-distribution
* license-installer
* metadata-editor
* schema-workbench
  {% endhint %}

4. Start PSW:

```bash
cd
cd ~/Pentaho/design-tools/schema-workbench
./workbench.sh
```

<figure><img src="/files/7ZnRs4QRFOFgqahp9bwk" alt="Pentaho Schema Workbench UI"><figcaption><p>Schema Workbench</p></figcaption></figure>
{% endtab %}

{% tab title="4. Aggregation Designer" %}
{% hint style="info" %}

#### **Aggregation Designer (PAD)**

Aggregation Designer recommends and builds aggregate tables to optimize Analyzer queries.
{% endhint %}

1. Extract `pad-ee-*.zip`.

```bash
cd ~/Pentaho/design-tools
unzip ~/Downloads/'Client Tools'/'Aggregation Designer'/pad-ee-10.2.0.0-222.zip
```

2. Make `.sh` files executable.

```bash
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec chmod +x {} \;
```

3. Verify structure:

{% hint style="info" %}
\~/Pentaho/design-tools/

* data-integration
* jdbc-distribution
* license-installer
* metadata-editor
* schema-workbench
* aggregation-designer
  {% endhint %}

4. Start PAD:

```bash
cd ~/Pentaho/design-tools/aggregation-designer
./startaggregationdesigner.sh
```

<figure><img src="/files/SyrwimuhfiWSzx0f33lI" alt="Pentaho Aggregation Designer UI"><figcaption><p>Aggregation Designer</p></figcaption></figure>
{% endtab %}
{% endtabs %}

<details>

<summary>General Troubleshooting (click to expand)</summary>

* Spoon fails to start or shows GTK/WebKit errors:
  * Install GTK/WebKit libs (see the note below) and try again.
  * Launch with additional SWT/GTK flags if needed, or test on a different desktop flavor.
* Fonts/UI rendering issues on HiDPI displays:
  * Try `GDK_SCALE=2 ./spoon.sh` or adjust your desktop scaling.
* Missing JDBC drivers in clients:
  * Copy the required driver JARs into the tool‑specific folders (see driver locations in the Server page) and restart the tool.
* License prompts or feature disabled:
  * Ensure the License Manager has activated client entitlements; verify `PENTAHO_LICENSE_INFORMATION_PATH` if required.
* Slow startup or out‑of‑memory errors:
  * Increase `-Xms`/`-Xmx` in the tool’s `*.ini` or startup script.

</details>

<details>

<summary>Missing GTK/WebKit modules</summary>

You will need to add a version from previous Jammy release:

1. Add package repository.

```bash
sudo apt-get install -qq software-properties-common
```

2. Add repository entry.

```bash
sudo apt-key adv --keyserver keyserver.ubuntu.com --recv-keys 3B4FE6ACC0B21F32
sudo add-apt-repository 'deb [trusted=yes] http://cz.archive.ubuntu.com/ubuntu bionic main universe'
```

3. Update repositories.

```bash
sudo apt-get update
```

5. Install package.

```bash
sudo apt-get install -qq libwebkitgtk-1.0-0
```

```bash
sudo apt-get install libcanberra-gtk-module
```

6. Start PDI.

```bash
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

</details>

***


# Kettle Plugins

Extend functionality with EE plugins ..

x

{% tabs %}
{% tab title="NEW - PDI Plugin Manager" %}
{% hint style="info" %}

#### PDI Plugin Manager

Pentaho Data Integration (PDI) can be extended with plugins that add new steps, job entries, and other functionality. The best way to manage these plugins is through the Plugin Manager, which you'll find in both the PDI client and Pentaho User Console (PUC).

The Plugin Manager handles all your plugin needs: installing new ones, updating existing ones to their latest versions, and removing plugins you no longer use.

While you can install plugins manually, this approach isn't recommended. Manually installed plugins won't show up in the Plugin Manager, which means you'll have to handle all future updates and removals yourself.
{% endhint %}

1. In the top toolbar Select: Tools > Plugin Manager.

<figure><img src="/files/ixVbzaZMbLR41vGHEimG" alt=""><figcaption><p>PDI Plugin Manager</p></figcaption></figure>

**Installing a Plugin:** Find the plugin you want to install by searching or browsing the available options.

**For the latest version:** Simply click Install.

<figure><img src="/files/tFatc4WOrBH0aDSQRGjs" alt=""><figcaption></figcaption></figure>

**For an earlier version:** Click on the plugin's table row to open the Plugin name dialog box. Select your desired version from the dropdown list and click Install. Confirm the installation if prompted.

**Restart to activate:** After installation, restart both Pentaho Server & PDI client. This step is essential - newly installed plugins won't work until you restart.

**Verify the installation:** Log into the PDI client and navigate to Tools > Plugin Manager. Search for or browse to your newly installed plugin. Check the Installed Version column to confirm the correct version is listed.
{% endtab %}

{% tab title="Plugin Matrix" %}

<table><thead><tr><th width="225">Plugin</th><th>Description</th></tr></thead><tbody><tr><td>Databricks</td><td>The Bulk load into Databricks entry loads large volumes of data from cloud storage files directly into Databricks tables. <strong>How it works:</strong> It accomplishes this by using Databricks' <a href="https://docs.databricks.com/aws/en/sql/language-manual/delta-copy-into">COPY INTO</a> command behind the scenes.</td></tr><tr><td>Salesforce Bulk Operation</td><td><p>The Salesforce bulk operation step performs large-scale data operations (insert, update, upsert, and delete) on Salesforce objects using the Salesforce Bulk API 2.0.</p><p><strong>How it works:</strong> The step reads data from an input stream, creates a CSV file of the changes, and executes the bulk job against Salesforce. After the job completes, you can optionally route three types of results to separate output streams: successful records, unprocessed records, and failed records.</p><p><strong>Requirements:</strong> You must have a Salesforce Client ID and Client Secret to use this step.</p></td></tr><tr><td>Google Analytics v4</td><td><p>The Google Analytics v4 step retrieves data from your Google Analytics account for reporting or data warehousing purposes.</p><p><strong>How it works:</strong> The step queries Google Analytics properties through the <a href="https://developers.google.com/analytics/devguides/reporting/data/v1">Google Analytics API v4</a> and sends the resulting dimension and metric values to the output stream.</p></td></tr><tr><td><a href="https://academy.pentaho.com/pentaho-data-integration/data-integration/ee-plugins/hierarchical-data-type">Hierarchical Data Type</a></td><td><p>Pentaho supports a hierarchical data type (HDT) through the Pentaho EE Marketplace plugin. This plugin adds the HDT data type and includes five specialized steps for working with it.</p><p><strong>What it does:</strong> These steps simplify working with complex, nested data structures. They can convert between HDT fields and formatted strings, and let you directly access or modify nested array indices and keys.</p><p><strong>Performance benefits:</strong> The steps significantly improve performance compared to handling hierarchical data as plain strings.</p><p><strong>Data structure:</strong> HDT can store nested or complex data built from objects and arrays, as well as single elements. It's compatible with any PDI step that processes hierarchical data.</p></td></tr><tr><td>Kafka Job</td><td></td></tr><tr><td></td><td></td></tr></tbody></table>
{% endtab %}
{% endtabs %}

x


# Post Installation Tasks

Hardening & performance ..

{% hint style="info" %}

#### **Post‑installation Hardening & Tuning**

Optional settings you can apply after installation to harden Tomcat/Pentaho and tune behaviour:
{% endhint %}

<details>

<summary>Hide Tomcat Server header</summary>

By default, Tomcat sends a `Server` header exposing version information. You can override it to reduce information leakage.

1. Edit the Tomcat connector in `server.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/conf/server.xml
```

2. Add or update the `server` attribute on the HTTP connector and (if used) AJP connector, then save.

```xml
<Connector port="8080" protocol="HTTP/1.1"
           connectionTimeout="20000"
           server=" "
           redirectPort="8443" />
```

3. Restart Pentaho Server.

```bash
sudo systemctl restart pentaho-server
```

</details>

<details>

<summary>Java Security Manager (deprecated/removed)</summary>

The legacy Java Security Manager is deprecated and not available on modern Java LTS versions (including Java 21). Do not use `-security` with Tomcat on Java 21. Prefer OS‑level hardening, least‑privilege users, network scoping, and container/AppArmor/SELinux policies as appropriate.

</details>

<details>

<summary>Change the web application context path</summary>

Change the context path if you do not want the application accessible at `/pentaho`.

1. Stop the Pentaho Server.

```bash
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Edit `context.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/META-INF/context.xml
```

3. Update the context path.

```xml
<Context path="/company" docBase="webapps/company/" />
```

4. Rename the webapp folder to match the new context name.

```bash
sudo mv /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho \
        /opt/pentaho/server/pentaho-server/tomcat/webapps/company
```

5. Update the redirect in `ROOT/index.jsp`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/ROOT/index.jsp
```

Change the meta refresh to:

```html
<meta http-equiv="refresh" content="0;URL=/company">
```

6. Update the server URL.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/server.properties
```

```
fully-qualified-server-url=http://localhost:8080/company/
```

7. Start the server and test.

```bash
sudo ./start-pentaho.sh
```

{% hint style="warning" %}
Upgrades may overwrite deployed webapps. Reapply customizations after upgrades, or use reverse proxy path mapping instead.
{% endhint %}

</details>

<details>

<summary>Change to HTTPs</summary>

Default port is 8080.

1. Stop the Pentaho Server.

```bash
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Change the connector port.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/conf/server.xml
```

```xml
<Connector URIEncoding="UTF-8"
      port="8443"
      protocol="org.apache.coyote.http11.Http11NioProtocol"
      maxThreads="150"
      SSLEnabled="true"
      scheme="https"
      secure="true"
      clientAuth="false"
      sslProtocol="TLS"
      keystoreType="PKCS12"
      keystoreFile="/opt/pentaho/pentaho-server/tomcat/ssl/keystore.p12"
      keystorePass="changeit"
    />
```

3. Update the server URL.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/server.properties
```

```
fully-qualified-server-url=http://localhost:8090/pentaho/
```

4. Start the server and verify.

```bash
sudo ./start-pentaho.sh
curl -I http://localhost:8090/pentaho/ | head -n 1
```

</details>

<details>

<summary>Change default HTTP port</summary>

Default port is 8080.

1. Stop the Pentaho Server.

```bash
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Change the connector port.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/conf/server.xml
```

```xml
<Connector URIEncoding="UTF-8"
           port="8090" protocol="HTTP/1.1"
           connectionTimeout="20000"
           redirectPort="8443"
           relaxedPathChars="[]|"
           relaxedQueryChars="^{}[]|&amp;"
           maxHttpHeaderSize="65536" />
```

3. Update the server URL.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/server.properties
```

```
fully-qualified-server-url=http://localhost:8090/pentaho/
```

4. Start the server and verify.

```bash
sudo ./start-pentaho.sh
curl -I http://localhost:8090/pentaho/ | head -n 1
```

</details>

<details>

<summary>Harden or disable the Tomcat shutdown port</summary>

By default Tomcat listens on a local shutdown port (8005) for the `SHUTDOWN` command.

* Disable the port by setting `port="-1"`, or
* Change both the port and the shutdown command to unpredictable values.

1. Edit the `<Server>` element in `server.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/conf/server.xml
```

Examples:

```xml
<Server port="-1" shutdown="SHUTDOWN">
```

or

```xml
<Server port="18005" shutdown="My$tr0ngShutCmd">
```

2. Restart Pentaho Server.

```bash
sudo systemctl restart pentaho-server
```

</details>

<details>

<summary>Custom error pages (404, 403, 500)</summary>

Define application‑level error pages to avoid exposing defaults.

1. Create an error page in your webapp.

```bash
sudo tee /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/error.jsp >/dev/null <<'EOF'
<html>
<head>
  <title>Error</title>
</head>
<body>
  <h1>Something went wrong</h1>
  <p>Please contact your administrator.</p>
</body>
</html>
EOF
```

2. Add error mappings in the webapp `web.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/WEB-INF/web.xml
```

```xml
<error-page>
  <error-code>404</error-code>
  <location>/error.jsp</location>
</error-page>
<error-page>
  <error-code>403</error-code>
  <location>/error.jsp</location>
</error-page>
<error-page>
  <error-code>500</error-code>
  <location>/error.jsp</location>
</error-page>
```

3. Restart the server and test.

</details>

<details>

<summary>Session timeout</summary>

Set a global session timeout for the application.

1. Edit the webapp `web.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/WEB-INF/web.xml
```

```xml
<session-config>
  <session-timeout>20</session-timeout>
</session-config>
```

</details>

<details>

<summary>Increase Karaf startup wait time</summary>

If server startup times out while Karaf installs features, increase the wait time.

1. Stop the server.

```bash
sudo systemctl stop pentaho-server
```

2. Edit `server.properties`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/server.properties
```

Uncomment or add:

```
# Time (ms) to wait for Karaf to install features before timing out
karafWaitForBoot=180000
```

3. Start the server.

```bash
sudo systemctl start pentaho-server
```

</details>

<details>

<summary>Remove sample data from the server</summary>

Remove evaluation samples before moving to production.

1. Stop the server.

```bash
sudo systemctl stop pentaho-server
```

2. Delete the `samples.zip` from default content (path may vary by version).

```bash
sudo rm -f /opt/pentaho/server/pentaho-server/pentaho-solutions/system/default-content/samples.zip || true
```

3. Edit the webapp `web.xml` and remove the HSQLDB sample definitions and the SystemStatusFilter (dev‑only).

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/WEB-INF/web.xml
```

Remove blocks similar to:

```xml
<context-param>
  <param-name>hsqldb-databases</param-name>
  <param-value>sampledata@../../data/hsqldb/sampledata</param-value>
</context-param>

<listener>
  <listener-class>org.pentaho.platform.web.http.context.HsqldbStartupListener</listener-class>
</listener>

<filter>
  <filter-name>SystemStatusFilter</filter-name>
  <filter-class>com.pentaho.ui.servlet.SystemStatusFilter</filter-class>
</filter>
```

4. Optionally remove the server `data/` directory if only sample content was used (verify your environment before deleting).

```bash
sudo rm -rf /opt/pentaho/server/pentaho-server/data || true
```

5. Start the server and remove sample folders via PUC (Browse Files → Public → Move to Trash).

```bash
sudo systemctl start pentaho-server
```

</details>

<details>

<summary>Hide Home perspective widgets</summary>

Hide Getting Started and other widgets from the PUC Home page.

1. Stop the server.

```bash
sudo systemctl stop pentaho-server
```

2. Edit the Home perspective configuration.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/mantle/home/properties/config.properties
```

Add or update:

```
disabled-widgets=getting-started,recents,favorites
```

3. Start the server and log in to verify.

```bash
sudo systemctl start pentaho-server
```

</details>

<details>

<summary>Turn off autocomplete on the login page (advanced)</summary>

Changing vendor JSPs may be overwritten on upgrade. Prefer SSO or reverse proxy controls. If you must, edit the login JSP.

1. Stop the server.

```bash
sudo systemctl stop pentaho-server
```

2. Edit `PUCLogin.jsp`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/jsp/PUCLogin.jsp
```

3. Set autocomplete to off for user/password inputs.

```html
<input id="j_username" name="j_username" type="text" autocomplete="off">
<input id="j_password" name="j_password" type="password" autocomplete="off">
```

4. Start the server.

```bash
sudo systemctl start pentaho-server
```

</details>

<details>

<summary>Increase CSV upload limits</summary>

Adjust upload limits and (optionally) staging database.

1. Edit `pentaho.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/pentaho.xml
```

```xml
<file-upload-defaults>
  <relative-path>/system/metadata/csvfiles/</relative-path>
  <max-file-limit>10000000</max-file-limit>
  <max-folder-limit>500000000</max-folder-limit>
</file-upload-defaults>
```

2. Change the staging database for CSV files (optional) in `data-access/settings.xml`.

```bash
sudo nano /opt/pentaho/server/pentaho-server/pentaho-solutions/system/data-access/settings.xml
```

```xml
<!-- settings for Agile Data Access -->
<data-access-staging-jndi>hibernate</data-access-staging-jndi>
```

3. In PUC, go to Tools → Refresh System Settings, then restart PUC (or the server) to apply.

</details>

***


# Pentaho Upgrade & Patches

Pentaho Upgrade Installer ..

{% hint style="info" %}

#### **Pentaho Upgrade Installer**

You can upgrade your Pentaho products from version 8.3 or later to version 10.1 with the Pentaho Upgrade Installer.

The upgrade installer checks your environment for version 8.3 or later Pentaho products, creates a backup of these products, then upgrades them to version 10.1. The Pentaho Upgrade Installer works for any Pentaho products you have installed on your server or workstations, including your Pentaho Server and your Pentaho client tools.

The Pentaho Upgrade Installer requires 22 GB of free space to perform the upgrade process.
{% endhint %}

{% hint style="warning" %}
If you need to upgrade your Pentaho products from a version earlier than 8.3, such as 7.1 or higher, you must upgrade your products to version 8.3, then use the Pentaho Upgrade Installer to move from version 8.3 to 10.1.
{% endhint %}

{% tabs %}
{% tab title="Checklist" %}
{% hint style="info" %}
Before you can run the Pentaho Upgrade Installer, you must also perform the following tasks:
{% endhint %}

* [x] Verify that your system components are current.

{% embed url="<https://docs.pentaho.com/pdia-10.2-install/components-reference>" %}

* [x] If you are upgrading an environment that includes the Pentaho Server, stop the server prior to performing backups and installation.

```bash
cd 
cd ~/[Pentaho Installation Directory]/server/pentaho-server
sh stop-pentaho.sh
```

1. Review your customizations. During the upgrade process, you can help the upgrade installer specify which items contain your customizations. See [Specify customized items to address after upgrading](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/r2F5~x211KC0wwU0qcq6cQ) for details. Then, after upgrading your Pentaho products to 10.1, you can merge your previous customizations into post-upgrade versions of the Pentaho files. See the [Apply customizations](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/S2d88cUzhPmuc8jUpi9NaA) post-upgrade task for instructions.

{% hint style="warning" %}
The upgrade process does not retain the drivers for your Hadoop clusters. You will need to re-install your drivers after completing the upgrade process.
{% endhint %}

1. Note: The upgrade process does not retain the drivers for your Hadoop clusters. You will need to re-install your drivers after completing the upgrade process. See the [Install drivers for your Hadoop clusters](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/UEFngjwGGT~SZRKXjCWeNQ) post-upgrade task for details.
2. If you are using plugins with your Pentaho products, review and back up your plugins to a separate directory structure.

{% hint style="warning" %}
The upgrade process does not retain your plugins. You will need to re-apply your plugins after completing the upgrade process.
{% endhint %}

1.
2. See the [Apply your plugins](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/pKp_UIrYWFitwgefW2h9ig) post-upgrade task for details.
3. If you are upgrading the Pentaho Server, verify that no users are logged on to the server.As a best practice, perform the upgrade process of the Pentaho Server during off-business hours to minimize the impact on your day-to-day operations.
4. Before installing the Pentaho Upgrade, verify that you have the most recent version of Java installed and that the JAVA\_HOME environment variable is set to that version of Java.
   {% endtab %}

{% tab title="Release" %}
{% hint style="info" %}
You can upgrade your Pentaho products from version 8.3 or later to version 9.4 using the Pentaho Upgrade Installer.

The upgrade installer checks your environment for version 8.3 or later Pentaho products, creates a backup of these products, then upgrades them to version 9.4.

The Pentaho Upgrade Installer works for any Pentaho products you have installed on your server or workstations, including your Pentaho Server and your Pentaho client tools.
{% endhint %}

{% hint style="warning" %}
The Pentaho Upgrade Installer requires 22 GB of free space to perform the upgrade process.
{% endhint %}

Before you can run the Pentaho Upgrade Installer, you must also perform the following tasks:

* [ ] Verify that your system components are current.

{% embed url="<https://help.hitachivantara.com/Documentation/Pentaho/9.4/Setup/Components_Reference>" %}
Pentaho 9.4
{% endembed %}

* [ ] If you are upgrading an environment that includes the Pentaho Server, stop the server prior to performing backups and installation.

```bash
cd 
cd ~/Pentaho/server/pentaho-server
sh stop-pentaho.sh
```

* [ ] Review any customizations.

During the upgrade process, you can help the upgrade installer specify which items contain your customizations.
{% endtab %}
{% endtabs %}


# Windows Installation

Windows Evaluation Wizard ..

{% hint style="info" %}

#### Windows Installation

The Installation Wizard provides the easiest and quickest way to install Pentaho Enterprise on Windows.

In this workshop, you will install 'everything' enabling access to a 'local' Pentaho Repository.
{% endhint %}

1. Navigate to the downloaded: `pentaho-business-analytics-11.0.0.0-237.exe` installation file.
2. Double-click: `pentaho-business-analytics-11.0.0.0-237.exe` file to launch it.
3. Ignore the Antivirus warning ..

<figure><img src="/files/yp281gazmtvhnhrYCQgQ" alt=""><figcaption><p>Review installation</p></figcaption></figure>

4. Accept the License agreement.

<figure><img src="/files/2xrdyp4xlIq4oawl4TXA" alt=""><figcaption><p>Accept License Agreement</p></figcaption></figure>

5. Keep the default installation path.

<figure><img src="/files/w5vPKhnXdVNzcEicLqTl" alt=""><figcaption><p>Keep default installation path</p></figcaption></figure>

6. The default Pentaho Repository database is PostgreSQL.

Password: password (all in lower case)

<figure><img src="/files/3Y2jRVx36BMYGbIOlp3M" alt=""><figcaption><p>Enter Password: password</p></figcaption></figure>

7. Decide if you wish to install everything or specific components.

<figure><img src="/files/0Z1v8OvhahB0xBpbgkL7" alt=""><figcaption><p>Keep it simple</p></figcaption></figure>

{% hint style="info" %}
With the Pentaho Installation Wizard you can choose one of two ways to install Pentaho components:

* Default: Select the Keep it simple. Give me everything option in the installation wizard.
* Custom: Select the Let me decide for myself option in the installation wizard.
  {% endhint %}

<figure><img src="/files/Btqy81N3oYwDoXXqJZ5D" alt=""><figcaption></figcaption></figure>

8. Install Stell Wheels sample data.

<figure><img src="/files/nsMtMH9Q3mOOX8M53SmE" alt=""><figcaption><p>Install Steel Wheels - sampledata</p></figcaption></figure>

9. Enter either an Activation ID or a Licensing server URL.

{% hint style="warning" %}
If you leave it blank, a 30-day license will be installed.
{% endhint %}

<figure><img src="/files/PkjjhIqclh62DkSqV3sm" alt=""><figcaption><p>Enter License URL or Activation ID</p></figcaption></figure>

10. Click 'Next' to start the installation.

<figure><img src="/files/K0LM4x2WObqdCqgsoLu7" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Various notifications will appear during the installation process. Allow Pentaho services to be installed.
{% endhint %}

<figure><img src="/files/qjZEtJNEXfR3ksgp85xU" alt=""><figcaption><p>Pentaho 11</p></figcaption></figure>

{% hint style="danger" %}
Ensure the required Pentaho Server & Data Integration plugins are installed.
{% endhint %}

<figure><img src="/files/cLqABki0hJno0GGrNcnR" alt=""><figcaption><p>Data Integration - Plugin Manager</p></figcaption></figure>

<figure><img src="/files/OWVZTqeePwHVZJajV9Jy" alt=""><figcaption><p>Pentaho Server - Plugin Manager</p></figcaption></figure>

***


# Containers

{% hint style="info" %}

#### Containers

Containers are lightweight, standalone packages that include everything needed to run an application: code, runtime, system tools, libraries, and settings. Unlike virtual machines that virtualize hardware, containers virtualize the operating system, sharing the host OS kernel while isolating the application processes. This makes them much more efficient and faster to start.
{% endhint %}

<figure><img src="/files/tK3RgKmY00ZsVoL49PZk" alt=""><figcaption><p>Containers</p></figcaption></figure>

{% tabs %}
{% tab title="Runtimes" %}
{% hint style="info" %}
**Docker** is the most well-known container platform. It popularized containerization by making it accessible and easy to use. Docker provides tools for building container images (Dockerfile), running containers, and managing them. While "Docker" often refers to the entire platform, Docker Engine is the actual runtime that executes containers. Docker also includes Docker Compose for defining multi-container applications and Docker Swarm for basic orchestration.

**containerd** is an industry-standard core container runtime that actually powers Docker (Docker uses it under the hood). It manages the complete container lifecycle - image transfer, storage, execution, and networking. Many Kubernetes deployments use containerd directly rather than going through Docker.

**CRI-O** is a lightweight container runtime built specifically for Kubernetes. It implements the Kubernetes Container Runtime Interface (CRI) and is designed to be a minimal runtime for Kubernetes without extra features.

**Podman** is a daemonless container engine that's compatible with Docker commands but doesn't require a background service running with root privileges. It's popular in security-conscious environments and on Red Hat/Fedora systems.

**LXC/LXD** (Linux Containers) is an older container technology that provides OS-level virtualization. LXC containers are more like lightweight VMs, running a full Linux system, whereas Docker containers typically run single applications.
{% endhint %}

<figure><img src="/files/QAsnCsLMPpngzalDMF6c" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Orchestration" %}
{% hint style="info" %}
**Kubernetes** (K8s) is the dominant container orchestration platform. It doesn't run containers itself but manages container runtimes, automating deployment, scaling, networking, and management of containerized applications across clusters of machines. Kubernetes introduces concepts like pods (groups of containers), services, deployments, and namespaces to manage complex applications.

**Docker Swarm** is Docker's native orchestration tool, simpler than Kubernetes but less feature-rich. It's easier to set up but has largely been overshadowed by Kubernetes.

**Apache Mesos** with Marathon was an early container orchestration platform, though it's less common now. It can orchestrate both containers and other workloads.

**Nomad** by HashiCorp is a simpler alternative to Kubernetes that can orchestrate containers, VMs, and standalone applications.

**Amazon ECS/EKS, Azure Container Instances, Google Kubernetes Engine** are cloud-specific managed container services that handle much of the infrastructure complexity.
{% endhint %}

<figure><img src="/files/VamCIFYx61S0gX9WbMgu" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Specialized" %}
{% hint style="info" %}
**Windows Containers** allow containerization of Windows applications, though they're less common than Linux containers. They come in two types: Windows Server Containers (process isolation) and Hyper-V Containers (stronger isolation).

**Serverless Containers** like AWS Fargate, Azure Container Instances, and Google Cloud Run let you run containers without managing the underlying infrastructure - you just provide the container image.
{% endhint %}
{% endtab %}
{% endtabs %}

{% hint style="info" %}
The ecosystem has largely converged around OCI (Open Container Initiative) standards, meaning most tools are interoperable. Docker remains the most popular for development, containerd for production runtimes, and Kubernetes for orchestration at scale.
{% endhint %}


# Docker

Docker deployment - Pentaho 11 + PostgreSQL 15 repository ..

{% hint style="info" %}

#### Docker Container

Docker container deployment enables you to package and run Pentaho products within portable, production-ready containers. Containerization ensures consistent behavior across development, testing, and production environments while simplifying deployment and scaling operations.

You can create Docker containers for the Pentaho Server, which includes the complete Business Analytics and Data Integration platform with the Pentaho User Console, scheduling services, and repository management. The server supports enterprise database backends including PostgreSQL, MySQL, Oracle, and SQL Server.

For distributed ETL processing, you can deploy Carte server containers that execute transformations and jobs remotely. The Kitchen and Pan command-line tools are also available as containers, enabling integration with CI/CD pipelines and automated batch workflows.

Container deployments are particularly effective for cloud environments where you can quickly scale resources to match data processing demands. By running Pentaho workloads in containers, organizations can optimize infrastructure costs while maintaining the flexibility to move between on-premises and cloud platforms.
{% endhint %}

<figure><img src="/files/Bhejyo4bfBXKZLnXw0M7" alt=""><figcaption><p><em>Docker Container Architecture showing Pentaho Server and PostgreSQL containers</em></p></figcaption></figure>

{% tabs %}
{% tab title="Pentaho Server Container" %}
{% hint style="info" %}
This container runs the complete Pentaho Business Analytics and Data Integration platform on Apache Tomcat 10. It includes the Pentaho User Console (PUC), scheduling services, and all analytics capabilities.
{% endhint %}
{% endtab %}

{% tab title="PostgreSQL Container" %}
{% hint style="info" %}
This container provides the relational database backend using PostgreSQL 17. It hosts three critical databases required by Pentaho: Jackrabbit (content repository), Quartz (scheduler), and Hibernate (security and audit). Data is persisted through a Docker volume to survive container restarts.
{% endhint %}

<figure><img src="/files/64dertqWlxEri5oIxnDf" alt=""><figcaption><p><em>PostgreSQL Database Architecture</em></p></figcaption></figure>

Pentaho Server requires three separate databases, each serving a distinct purpose:

<table><thead><tr><th width="128" valign="top">Database</th><th width="133" valign="top">Owner</th><th valign="top">Purpose &#x26; Contents</th></tr></thead><tbody><tr><td valign="top">jackrabbit</td><td valign="top">jcr_user</td><td valign="top">Java Content Repository (JCR) - Stores all Pentaho content including reports, dashboards, data sources, analysis schemas, and user files. This is the primary content storage for the Pentaho repository.</td></tr><tr><td valign="top">quartz</td><td valign="top">pentaho_user</td><td valign="top">Quartz Scheduler - Manages all scheduled jobs, triggers, and calendars. Contains tables for job definitions (QRTZ6_JOB_DETAILS), triggers (QRTZ6_TRIGGERS), execution history, and cluster coordination locks.</td></tr><tr><td valign="top">hibernate</td><td valign="top">hibuser</td><td valign="top">Hibernate Repository - Hosts security configuration, audit logging, user session data, and contains two additional schemas: pentaho_dilogs (ETL execution logging) and pentaho_operations_mart (analytics data mart).</td></tr></tbody></table>

{% hint style="info" %}
The hibernate database contains specialized schemas for operational monitoring:

**pentaho\_dilogs**: Captures detailed ETL execution information including job logs, transformation logs, step performance metrics, and error records. Essential for debugging data integration workflows and monitoring pipeline health.

**pentaho\_operations\_mart**: A dimensional data mart for analytics on Pentaho usage. Contains dimension tables (DIM\_DATE, DIM\_TIME, DIM\_EXECUTOR) and fact tables (FACT\_EXECUTION, FACT\_STEP\_EXECUTION) for analyzing platform utilization, performance trends, and user activity.

For production deployments, implement regular backups of the repository-data Docker volume. The jackrabbit database is the most critical as it contains all user content. Consider using pg\_dump for logical backups or volume snapshots for full recovery options.
{% endhint %}
{% endtab %}

{% tab title="Network" %}
{% hint style="info" %}
Key points:

* The Pentaho container connects to PostgreSQL using the service name 'repository' as the hostname
* PostgreSQL listens on port 5432 internally (not exposed to host by default)
* Pentaho Server exposes port 8080, mapped to the host system
* All inter-container traffic remains within the Docker network for security
  {% endhint %}

<figure><img src="/files/O24xb2E97EdrynJByGhj" alt=""><figcaption><p><em>Data flow showing HTTP requests and JDBC connections</em></p></figcaption></figure>

Tomcat manages connection pools defined in context.xml. Each pool serves a specific purpose:

<table><thead><tr><th valign="top">Pool Name</th><th valign="top">Connection Target</th><th valign="top">Used For</th></tr></thead><tbody><tr><td valign="top">jdbc/Hibernate</td><td valign="top">repository:5432/hibernate</td><td valign="top">Security, Users, Roles</td></tr><tr><td valign="top">jdbc/Quartz</td><td valign="top">repository:5432/quartz</td><td valign="top">Job Scheduling</td></tr><tr><td valign="top">jdbc/jackrabbit</td><td valign="top">repository:5432/jackrabbit</td><td valign="top">Content Repository</td></tr><tr><td valign="top">jdbc/Audit</td><td valign="top">repository:5432/hibernate</td><td valign="top">Audit Logging</td></tr><tr><td valign="top">jdbc/live_logging_info</td><td valign="top">repository:5432/hibernate</td><td valign="top">ETL Runtime Logs</td></tr><tr><td valign="top">jdbc/PDI_Operations_Mart</td><td valign="top">repository:5432/hibernate</td><td valign="top">Operations Analytics</td></tr></tbody></table>

{% hint style="info" %}
When a user accesses Pentaho Server:

1\. User's browser sends HTTP request to localhost:8080

2\. Docker forwards the request to Pentaho container's port 8080

3\. Tomcat receives request and routes to pentaho.war web application

4\. Application retrieves/stores data via JDBC connection pools

5\. JDBC connections route to 'repository:5432' (PostgreSQL container)

6\. Response flows back through the same path to user's browser
{% endhint %}
{% endtab %}

{% tab title="Volume Mapping" %}
{% hint style="info" %}

#### Volume Mapping

The deployment uses both named Docker volumes and bind mounts for persistence and configuration:
{% endhint %}

```
Docker Volumes:
  vault_data             -> /vault/data
  pentaho_postgres_data  -> /var/lib/postgresql/data
  pentaho_solutions      -> /opt/pentaho/pentaho-server/pentaho-solutions
  pentaho_data           -> /opt/pentaho/pentaho-server/data
 
Bind Mounts:
  ./softwareOverride     -> /docker-entrypoint-init (ro)
  ./db_init_postgres     -> /docker-entrypoint-initdb.d (ro)
  ./postgres-config      -> /etc/postgresql/conf.d (ro)
  ./config/.kettle       -> /home/pentaho/.kettle
  ./config/.pentaho      -> /home/pentaho/.pentaho
  ./vault/config         -> /vault/config (ro)
  ./scripts              -> /scripts (ro)
```

{% endtab %}
{% endtabs %}

***

{% hint style="danger" %}
Before you begin the Docker deployment, ensure you have completed the Setup: [Pentaho Containers](/pentaho-11-installation-en/setup/pentaho-containers)
{% endhint %}

Run through the following steps to deploy Pentaho Server with PostgreSQL 15 repository.

{% tabs %}
{% tab title="1. Prepare Environment" %}
{% hint style="info" %}

#### Prepare Environment

Check Docker is up and running:

* Copy Pentaho-Server-PostgreSQL assets
* Copy over `pentaho-server-ee-11.0.0.0-237.zip`
* Verify Docker & Docker Compose
* Check ports
  {% endhint %}

{% hint style="danger" %}
Ensure you have downloaded: `pentaho-server-ee-11.0.0.0-237.zip`
{% endhint %}

1. Create project directory & copy over assets.

```bash
cd
cp -r ~/Workshop--Installation/Pentaho-Containers/On-Prem/Pentaho-Server-PostgreSQL .
```

2. Copy over the pentaho-server-ee-11.0.0.0-237.zip /docker/stagedArtefacts directory.

{% hint style="info" %}
If you have deployed an Archive Pentaho Server then copy from:

`/opt/pentaho/software/pentaho-server-ee-version`

Otherwise download package from the [Pentaho Customer Portal](https://support.pentaho.com/hc/en-us).
{% endhint %}

```bash
cd
cd ~/Pentaho-Server-PostgreSQL/docker/stagedArtifacts
cp /opt/pentaho/software/server/pentaho-server-ee-11.0.0.0-237.zip . 
```

3. Verify that the file.

```bash
cd
cd ~/Pentaho-Server-PostgreSQL/docker/stagedArtifacts
ls -al
```

4. Check the Docker version.

```bash
docker --version
# Expected output: Docker version 29.0.2 or higher
```

5. Check Docker Compose version.

```bash
docker compose --version
# Expected output: Docker Compose version 2.40.3 or higher
```

```bash
sudo apt install docker-compose
# installs Docker Compose
```

6. Verify Docker daemon is running.

```bash
docker info
# Should display system-wide information without errors
```

7. Check port 8080 / 8090 is available on Host OS.

```bash
sudo lsof -i ::8080
```

{% hint style="info" %}
If port 8080 is in use by another application, you can change the PORT variable in the .env file to any available port (e.g. 8090, 8081, 9090).
{% endhint %}

8. Pentaho Server requires a valid license. The `.env` file contains a LICENSE\_URL pointing to the Flexera license server. Ensure your license entitlements are active before deployment.

{% hint style="warning" %}
Without a valid license, Pentaho Server will start but many features will be disabled. Verify your license status before proceeding with production deployments.
{% endhint %}
{% endtab %}

{% tab title="2. Directory Layout" %}
{% hint style="info" %}

#### Directory Layout

This deployment configuration provides several important capabilities:

* Completely self-contained and portable deployment.
* Automated database initialization with SQL scripts.
* Health checks and proper startup ordering between services.
* Persistent data volumes for database and Pentaho content.
* HashiCorp Vault for secrets management with AppRole authentication.
* Read-only containers with tmpfs mounts for security.
* Resource limits (CPU/memory) for stability.
* Log rotation to prevent disk exhaustion.
* Software override system for customizing configurations without modifying core files.
* Production-ready configuration templates.
* PostgreSQL JDBC driver included
* Easy backup and restore procedures
  {% endhint %}

{% hint style="success" %}
Check out the other repository deployment options at:

```
~/Workshop--Installation/Pentaho-Containers/On-Prem/
```

{% endhint %}

***

**Root Directory Files**

```
Pentaho-Server-PostgreSQL/
├── README.md             # Main documentation file
├── ARCHITECTURE.md       # System architecture details
├── CONFIGURATION.md      # Configuration reference guide
├── TROUBLESHOOTING.md    # Problem solving guide
├── docker-compose.yml    # Docker Compose service definitions
├── Makefile              # Convenience targets (make help)
├── deploy.sh             # Automated deployment script
├── .env                  # Environment configuration (created)
├── .env.template         # Environment template with defaults
```

{% hint style="info" %}
**Documentation Files:**

**README.md** - The main entry point documentation providing project overview, quick start instructions, prerequisites, and general usage information for the workshop.

**ARCHITECTURE.md** - Detailed technical documentation covering the system architecture, component relationships, container design, networking, data flow, and architectural decisions for the Docker-based deployment.

**CONFIGURATION.md** - Comprehensive configuration reference guide detailing all available environment variables, configuration options, customization parameters, and settings for both Pentaho Server and PostgreSQL components.

**TROUBLESHOOTING.md** - Problem-solving guide with common issues, error messages, diagnostic procedures, and solutions for deployment and runtime problems you might encounter.

**Orchestration & Deployment:**

**docker-compose.yml** - The Docker Compose service definitions file that declares all containers (Pentaho Server, PostgreSQL, potentially Vault/other services), their configurations, networking, volumes, and dependencies.

**Makefile** - Contains convenience command targets for common operations like building, starting, stopping, and cleaning up the environment. Users can run `make help` to see available commands.

**deploy.sh** - Automated deployment script that likely handles the complete deployment workflow including environment validation, building images, starting services, and initial configuration.

**Environment Configuration:**

**.env** - The active environment configuration file (created from template) containing actual values for database passwords, ports, hostnames, and other environment-specific settings. This file is typically git-ignored.

**.env.template** - The template file with default values and placeholders that users copy to create their `.env` file, providing a reference for all configurable environment variables.
{% endhint %}

***

**Docker Build Context**

```
├── docker/
│   ├── Dockerfile                # Multi-stage Pentaho Server image build
│   ├── entrypoint/
│   │   └── docker-entrypoint.sh  # Container startup script
│   └── stagedArtifacts/
│       └── pentaho-server-ee-11.0.0.0-237.zip
```

{% hint style="info" %}
The **docker/** directory contains all the core components needed to build and run the Pentaho Server containerized deployment:

**Dockerfile** - This is the main build configuration using a multi-stage build approach to create the Pentaho Server container image. Multi-stage builds help optimize the final image size by separating the build environment from the runtime environment.

**entrypoint/** directory contains the **docker-entrypoint.sh** script, which is the initialization script that runs when the container starts. This typically handles tasks like environment setup, configuration management, health checks, and starting the Pentaho Server services.

**stagedArtifacts/** directory serves as the staging area for the Pentaho Server installation package. It currently contains **pentaho-server-ee-11.0.0.0-237.zip**, which is the Enterprise Edition version 11.0.0.0 build 237 that gets extracted and installed during the Docker image build process.
{% endhint %}

***

**PostgreSQL Repsoitory Database Initialization**

```
├── db_init_postgres/
│   ├── 1_create_jcr_postgresql.sql         # Jackrabbit content repo
│   ├── 2_create_quartz_postgresql.sql      # Quartz scheduler
│   ├── 3_create_repository_postgresql.sql  # Hibernate repository
│   ├── 4_pentaho_logging_postgresql.sql    # Audit/DI logging schema
│   └── 5_pentaho_mart_postgresql.sql       # Operations mart schema
```

{% hint style="info" %}
The **db\_init\_postgres/** directory contains the PostgreSQL database initialization scripts that set up all the required schemas for Pentaho Server 11. These scripts are numbered to execute in a specific sequence:

**1\_create\_jcr\_postgresql.sql** - Creates the **Jackrabbit Content Repository (JCR)** schema, which stores the Pentaho repository content including solution files, schedules, reports, dashboards, and metadata. This is the core content management system for Pentaho.

**2\_create\_quartz\_postgresql.sql** - Sets up the **Quartz Scheduler** schema, which manages all scheduled jobs and tasks within Pentaho Server, including report generation, ETL executions, and other automated processes.

**3\_create\_repository\_postgresql.sql** - Creates the **Hibernate Repository** schema, which stores user authentication, authorization data, roles, permissions, and other security-related information managed by Pentaho's security subsystem.

**4\_pentaho\_logging\_postgresql.sql** - Establishes the **Audit and Data Integration (DI) Logging** schema for capturing execution logs, transformation/job metrics, and audit trail information from PDI processes running on the server.

**5\_pentaho\_mart\_postgresql.sql** - Creates the **Operations Mart** schema, which stores operational analytics data about Pentaho Server usage, performance metrics, and system monitoring information used by the Pentaho Operations Mart dashboard.
{% endhint %}

***

**PostgreSQL and Vault Configuration**

```
├── postgres-config/
│   ├── custom.conf             # PostgreSQL performance tuning
│   └── pg_hba.conf             # Client authentication config
├── vault/
│   ├── config/
│   │   └── vault.hcl           # Vault server configuration
│   └── policies/
│       └── pentaho-policy.hcl  # Pentaho access policy
├── secrets/
│   └── postgres_password.txt   # Docker secrets file
```

{% hint style="info" %}
**postgres-config/** - PostgreSQL Configuration

**custom.conf** - Custom PostgreSQL performance tuning parameters optimized for Pentaho Server workloads. This likely includes settings for shared buffers, work memory, connection limits, checkpoint configurations, and other performance-related parameters tailored to handle Pentaho's database requirements.

**pg\_hba.conf** - PostgreSQL Host-Based Authentication configuration file that controls client connection authentication methods, IP address access rules, and security policies for database connections from the Pentaho Server container.

***

**vault/** - HashiCorp Vault Integration

This directory supports incorporating Vault for secrets management:

**config/vault.hcl** - The HashiCorp Vault server configuration file defining storage backend, listener settings, API endpoints, seal/unseal behavior, and general Vault server operational parameters.

**policies/pentaho-policy.hcl** - Vault access control policy specifically for Pentaho Server, defining which secrets paths the Pentaho application can read, write, or manage. This enforces least-privilege access to sensitive credentials.

***

**secrets/** - Docker Secrets Management

**postgres\_password.txt** - A Docker secrets file containing the PostgreSQL password. When using Docker secrets (or Vault integration), this file provides the database credentials in a secure manner rather than passing them as plain environment variables. The file should have restricted permissions and is typically referenced by Docker Compose using the `secrets:` configuration.
{% endhint %}

<table><thead><tr><th valign="top">Key</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top">postgres_password</td><td valign="top">PostgreSQL superuser password</td></tr><tr><td valign="top">pentaho_user</td><td valign="top">Pentaho database username</td></tr><tr><td valign="top">pentaho_password</td><td valign="top">Pentaho database password</td></tr><tr><td valign="top">jdbc_url</td><td valign="top">JDBC connection URL</td></tr></tbody></table>

***

**Pentaho Configuration Overrides**

```
├── softwareOverride/
│   ├── 1_drivers/                              # JDBC drivers
│   │   └── tomcat/lib/
│   ├── 2_repository/                           # Database configuration
│   │   ├── pentaho-solutions/system/
│   │   │   ├── hibernate/
│   │   │   ├── jackrabbit/
│   │   │   └── scheduler-plugin/quartz/
│   │   └── tomcat/webapps/pentaho/META-INF/
│   ├── 3_security/                             # Authentication settings
│   │   └── pentaho-solutions/system/
│   └── 4_others/                               # Tomcat and app settings
│       ├── pentaho-solutions/system/
│       └── tomcat/
```

{% hint style="info" %}
The **softwareOverride/** directory contains customized configuration files and components that override the default Pentaho Server installation. The numbered structure ensures a logical organization and potentially an ordered application during the Docker build process:

***

**1\_drivers/** - JDBC Database Drivers

**tomcat/lib/** - Contains JDBC driver JAR files (specifically the PostgreSQL JDBC driver) that get copied into Tomcat's library directory, enabling Pentaho Server to connect to PostgreSQL databases.

***

**2\_repository/** - Database Repository Configuration

This section configures Pentaho's connection to all PostgreSQL-backed repositories:

**pentaho-solutions/system/hibernate/** - Hibernate repository configuration files (repository.xml, hibernate-settings.xml) for user/role security data

**pentaho-solutions/system/jackrabbit/** - Jackrabbit JCR repository configuration (repository.xml) for content storage

**pentaho-solutions/system/scheduler-plugin/quartz/** - Quartz scheduler database configuration (quartz.properties) for job scheduling

**tomcat/webapps/pentaho/META-INF/** - Contains context.xml with JNDI datasource definitions for all Pentaho databases (Quartz, Jackrabbit, Hibernate, Audit, Operations Mart)

***

**3\_security/** - Authentication & Security Settings

**pentaho-solutions/system/** - Security configuration files including applicationContext-security.xml, security.properties, and potentially LDAP/SSO configurations for authentication and authorization.

***

**4\_others/** - Additional Tomcat & Application Settings

**pentaho-solutions/system/** - Other system-level configurations like pentaho.xml, pentaho-spring-beans.xml, log4j settings, and application behavior configurations

**tomcat/** - Tomcat server customizations including server.xml, web.xml, setenv.sh for JVM parameters, and other Tomcat-specific tuning
{% endhint %}

**Utility Scripts**

```
├── scripts/
│   ├── backup-postgres.sh      # Database backup utility
│   ├── restore-postgres.sh     # Database restore utility
│   ├── backup-vault.sh         # Vault credentials backup
│   ├── restore-vault.sh        # Vault credentials restore
│   ├── rotate-secrets.sh       # Password rotation script
│   ├── fetch-secrets.sh        # Secret retrieval helper
│   ├── vault-init.sh           # Vault initialization
│   └── validate-deployment.sh  # Deployment validation
```

{% hint style="info" %}
The **scripts/** directory contains operational and maintenance utilities for managing the Pentaho Server deployment, organized by functional area:

***

**Database Management:**

**backup-postgres.sh** - Automated PostgreSQL backup utility that creates dumps of all Pentaho databases (JCR, Quartz, Hibernate, Audit, Operations Mart). Likely includes timestamping, compression, and backup retention logic.

**restore-postgres.sh** - Database restoration utility to recover Pentaho databases from backup files, useful for disaster recovery, environment cloning, or migrating data between instances.

***

**Vault/Secrets Management:**

**backup-vault.sh** - HashiCorp Vault credentials and unseal keys backup script, ensuring recovery capability for the Vault instance containing sensitive Pentaho credentials.

**restore-vault.sh** - Vault restoration utility to recover Vault data and re-initialize the secrets management system from backup.

**rotate-secrets.sh** - Automated password rotation script that updates database passwords and other sensitive credentials in Vault, then propagates changes to Pentaho Server configuration - supporting security best practices.

**fetch-secrets.sh** - Helper utility to retrieve secrets from Vault programmatically, useful for scripts that need to access credentials without hardcoding them.

**vault-init.sh** - Initial Vault setup script that handles Vault initialization, unsealing, creating the Pentaho policy, and storing initial secrets for the deployment.

***

**Operations & Validation:**

**validate-deployment.sh** - Deployment validation script that performs health checks on all components (PostgreSQL connectivity, Pentaho Server startup, Vault accessibility, service availability), confirming the environment is properly configured and operational.
{% endhint %}

***

**User Configuration and Data Storage**

```
├── config/
│   ├── .kettle/                # PDI/Kettle configuration
│   │   └── kettle.properties
│   └── .pentaho/               # Pentaho user settings
├── backups/                    # Database backup storage
│   └── *.sql.gz                # Compressed SQL backups
└── logs/                       # Application logs (optional)
```

{% hint style="info" %}
**config/** - Application Configuration

This directory stores user-level and application-level configuration files:

**`.kettle/`** - PDI (Pentaho Data Integration) / Kettle configuration directory

* **kettle.properties** - Contains Kettle/PDI environment variables, connection parameters, system properties, and global settings used by transformation and job executions running on Pentaho Server.

**`.pentaho/`** - Pentaho user settings directory for storing user-specific preferences, cached metadata, and application state information.

***

**backups/** - Database Backup Storage

**`*.sql.gz`** - Repository for compressed PostgreSQL database backup files created by the `backup-postgres.sh` script. The gzip compression reduces storage requirements while maintaining complete database snapshots for disaster recovery, environment cloning, or rollback scenarios. Backup files are likely timestamped for version tracking.

***

**logs/** - Application Logging

Centralized logging directory for capturing runtime logs from all services. This likely includes:

* Pentaho Server application logs (catalina.out, pentaho.log)
* PostgreSQL database logs
* Vault service logs
* Docker container logs
* ETL execution logs

This supports the **log rotation configuration** and monitoring capabilities you've been incorporating into your deployment, making troubleshooting and auditing easier during workshops.
{% endhint %}

***

**Key Files**

<table><thead><tr><th valign="top">File</th><th valign="top">Purpose</th></tr></thead><tbody><tr><td valign="top">docker-compose.yml</td><td valign="top">Defines all services (pentaho-server, postgres), networks, and volumes</td></tr><tr><td valign="top">docker/Dockerfile</td><td valign="top">Multi-stage build using debian:trixie-slim with OpenJDK 21</td></tr><tr><td valign="top">docker-entrypoint.sh</td><td valign="top">Processes softwareOverride directories at container startup</td></tr><tr><td valign="top">.env</td><td valign="top">Environment-specific configuration (ports, passwords, memory)</td></tr><tr><td valign="top">deploy.sh</td><td valign="top">Automated deployment with pre-flight validation checks</td></tr><tr><td valign="top">db_init_postgres/*.sql</td><td valign="top">PostgreSQL database initialization scripts</td></tr><tr><td valign="top">vault-init.sh</td><td valign="top">Initializes Vault and stores secrets</td></tr><tr><td valign="top">rotate-secrets.sh</td><td valign="top">Rotates database passwords securely</td></tr></tbody></table>
{% endtab %}

{% tab title="3. Pre-flight Tasks" %}
{% hint style="info" %}

#### Pre-flight Taks

The Pre-flight Tasks section outlines the essential preparation steps needed before deploying Pentaho Server 11 in Docker containers.

First, you need to configure the environment variables by editing the `.env.template` file with your deployment-specific settings. This includes defining the Pentaho version and image details, PostgreSQL credentials and port configuration (defaulting to 5432), Pentaho HTTP and HTTPS ports (8090 and 8443), JVM memory allocation (minimum 4GB, maximum 8GB), the license server URL, and Vault port settings. Once configured, this template is saved as the active `.env` file.

PostgreSQL performance tuning is handled through the `postgres-config/custom.conf` file, where you can customize connection limits (defaulting to 200 max connections), memory allocation parameters including shared buffers and cache sizes, and other performance optimizations specifically tuned for containerized environments.

Finally, the `softwareOverride/` directory provides an optional mechanism for customizing Pentaho configurations without modifying core installation files. The PostgreSQL JDBC driver comes included by default, but you can optionally upgrade it by downloading from Maven Central or copying from the workshop's database drivers collection. This preparation ensures all required files, configurations, and credentials are properly staged before running the automated deployment script.
{% endhint %}

***

**Configure .env**

1. Edit the .env.template

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
nano .env.template
```

2. Enter the following details:

<table><thead><tr><th valign="top">Variable</th><th valign="top">Default</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top">PENTAHO_VERSION</td><td valign="top">11.0.0.0-237</td><td valign="top">Pentaho Server version</td></tr><tr><td valign="top">PENTAHO_IMAGE_NAME</td><td valign="top">pentaho/pentaho-server</td><td valign="top">Docker image name</td></tr><tr><td valign="top">PENTAHO_IMAGE_TAG</td><td valign="top">11.0.0.0-237</td><td valign="top">Docker image tag</td></tr><tr><td valign="top">POSTGRES_PASSWORD</td><td valign="top">password</td><td valign="top">PostgreSQL root password</td></tr><tr><td valign="top">POSTGRES_PORT</td><td valign="top">5432</td><td valign="top">PostgreSQL exposed port</td></tr><tr><td valign="top">PENTAHO_HTTP_PORT</td><td valign="top">8090</td><td valign="top">Pentaho HTTP port</td></tr><tr><td valign="top">PENTAHO_HTTPS_PORT</td><td valign="top">8443</td><td valign="top">Pentaho HTTPS port</td></tr><tr><td valign="top">PENTAHO_MIN_MEMORY</td><td valign="top">4096m</td><td valign="top">JVM minimum heap size</td></tr><tr><td valign="top">PENTAHO_MAX_MEMORY</td><td valign="top">8192m</td><td valign="top">JVM maximum heap size</td></tr><tr><td valign="top">LICENSE_URL</td><td valign="top">(empty)</td><td valign="top">EE license server URL</td></tr><tr><td valign="top">VAULT_PORT</td><td valign="top">8200</td><td valign="top">Vault API port</td></tr></tbody></table>

3. Save:

```
CTRL + o
Enter
CTRL + x
```

4. Create .env

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
cp .env.template .env
```

***

**Customize postgres-config/custom.conf**

1. Edit the .env.template

```bash
cd
cd ~/Pentaho-Server-PostgreSQL/progres-config
nano custom.conf
```

2. Enter the following details:

```conf
# Connection limits
max_connections = 200

# Memory (adjust based on available RAM)
shared_buffers = 256MB
effective_cache_size = 768MB
work_mem = 16MB

# Performance
random_page_cost = 1.1
effective_io_concurrency = 200
```

3. Save:

```
CTRL + o
Enter
CTRL + x
```

***

**softwareOverride**

{% hint style="info" %}
The `softwareOverride/` directory provides a powerful mechanism to customize Pentaho Server without modifying the core installation. Files are copied into the Pentaho installation during container startup, processed in alphabetical order by directory name.
{% endhint %}

````
```
softwareOverride/
├── 1_drivers/           # JDBC drivers and data connectors
│   ├── tomcat/lib/
│   │   └── postgresql-42.x.x.jar    # PostgreSQL JDBC driver (included)
│   └── pentaho-solutions/drivers/    # Big data drivers (.kar files)
├── 2_repository/        # Database repository configuration
│   ├── pentaho-solutions/system/
│   │   ├── hibernate/hibernate-settings.xml
│   │   ├── jackrabbit/repository.xml
│   │   └── scheduler-plugin/quartz/quartz.properties
│   └── tomcat/webapps/pentaho/META-INF/context.xml
├── 3_security/          # Authentication and authorization
│   └── pentaho-solutions/system/
│       ├── applicationContext-spring-security-hibernate.properties
│       └── applicationContext-spring-security-memory.xml
├── 4_others/            # Tomcat, defaults, and miscellaneous
│   ├── pentaho-solutions/system/
│   │   ├── defaultUser.spring.properties
│   │   ├── pentaho.xml
│   │   └── security.properties
│   └── tomcat/
│       ├── bin/startup.sh
│       └── webapps/pentaho/WEB-INF/web.xml
└── 99_exchange/         # User data exchange (not auto-processed)
```
````

The PostgreSQL JDBC driver is included in the Pentaho distribution. If you need to upgrade:

1. Download from [Maven Central](https://repo1.maven.org/maven2/org/postgresql/postgresql/)
2. Place in `softwareOverride/1_drivers/tomcat/lib/`

Or

Copy from Workshop--Installation/'Database Drivers'/

```bash
cd
cd ~/Workshop--Installation/'Database Drivers'
cp postgresql-42.7.8.jar ~/Pentaho-Server-PostgreSQL/softwareOverride/1_drivers/tomcat/lib
```

{% endtab %}

{% tab title="4. Deployment" %}
{% hint style="info" %}

#### Deployment

This section walks through the deployment process using either the automated script or manual commands.
{% endhint %}

Select Deployment option:

{% tabs %}
{% tab title="Automated" %}
{% hint style="info" %}

#### Automated Deployment

The deploy.sh script automates the entire deployment process with pre-flight validation:

./deploy.sh

The script performs the following actions:

* Validates Docker and Docker Compose installation
* Verifies Pentaho package exists in: `docker/stagedArtifacts/`
* Creates `.env` from template if missing
* Checks disk space (10GB minimum)
* Verifies required ports are available
* Builds the Pentaho Server Docker image
* Starts PostgreSQL and waits for health check
* Starts Pentaho Server and monitors startup
* Displays access URLs and credentials
  {% endhint %}

1\. Set execute permissions on the deployment scripts.

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
chmod +x deploy.sh
chmod +x scripts/*.sh
```

2. Deploy the containers.

{% hint style="danger" %}
Ensure you dont have a postgresql service up and running:

```bash
systemctl stop postgresql
```

{% endhint %}

```bash
cd
cd ~/Pentaho-Server-PostgreSQL && ./deploy.sh
```

{% tabs %}
{% tab title="Pre-flight & Building Phase" %}
{% hint style="info" %}
**Pre-Flight Checks ✓**

The script validates the environment before starting:

* Docker is installed
* Docker Compose is installed
* Docker daemon is running
* Pentaho package found
* `.env` file exists
* Sufficient disk space (414GB available)
* Port 8090 (Pentaho HTTP) is available
* Port 5432 (PostgreSQL) is available

**Building Phase**

A custom Docker image is built with **24 build steps** taking approximately **5-10 minutes**:

* Base image: `debian:trixie-slim`
* Installs system packages via `apt-get update` and `apt-get upgrade`
* Installs **OpenJDK 21 JRE headless** with `curl` and `rm`
* Creates a `pentaho` user and group (GID 5000)
* (Optional) Installs Pentaho plugins (PAZ, PIR, PDD)
* Copies Pentaho installation to `/opt/pentaho/`
* Exports layers and manifests
* **Final image**: `pentaho/pentaho-server:11.0.0.0-237`
  {% endhint %}

<figure><img src="/files/DHr0ULb81axeVfO3xtLT" alt=""><figcaption><p>Pre-flight checks &#x26; Build</p></figcaption></figure>
{% endtab %}

{% tab title="Deployment & Final Status" %}
{% hint style="info" %}
**Starting PostgreSQL Database**

Pulls **PostgreSQL 15** image and related layers:

* Creates network: `pentaho-server-postgresql_pentaho-net`
* Creates volume: `pentaho-server-postgresql_pentaho_postgres_data`
* Creates container: `pentaho-postgres`
* Waits for PostgreSQL readiness: **✓ PostgreSQL is ready**

**Starting Pentaho Server**

This phase takes **2-3 minutes for first-time initialization**:

**Pulls HashiCorp Vault 1.15 image** for secrets management

**Creates volumes:**

* `pentaho-server-postgresql_pentaho_solutions`
* `pentaho-server-postgresql_pentaho_data`
* `pentaho-server-postgresql_vault_data`
  {% endhint %}

<figure><img src="/files/W6Bldy0xgCA17N3X7Ka8" alt=""><figcaption><p>Deploy Containers</p></figcaption></figure>

{% hint style="info" %}
**Final Status** 🎉

The deployment is successful and provides you with:

**Pentaho Server Access:**

* URL: `http://localhost:8090/pentaho`
* Login: `admin` / `password`

**PostgreSQL Database:**

* Host: `localhost:5432`
* Login: `postgres` / `password`
  {% endhint %}

| Action           | Command                  |
| ---------------- | ------------------------ |
| View logs        | `docker compose logs -f` |
| Stop services    | `docker compose stop`    |
| Start services   | `docker compose start`   |
| Restart services | `docker compose restart` |
| Shutdown         | `docker compose down`    |

{% hint style="info" %}
**Helper scripts provided:**

* `./scripts/backup-postgres.sh` -- Backup database
* `./scripts/restore-postgres.sh <backup-file>` -- Restore database
* `./scripts/validate-deployment.sh` -- Validate deployment
  {% endhint %}

<figure><img src="/files/2SSnhsbBv6OkmRCQQkjF" alt=""><figcaption><p>Helper</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Manual Build" %}
{% hint style="info" %}

#### Manual Build

{% endhint %}

1. Build the Pentaho Server image.

```bash
docker compose build --no-cache pentaho-server
```

{% hint style="info" %}
This process takes approximately 5-10 minutes as it extracts the Pentaho package and configures the image.
{% endhint %}

2. Start PostgreSQL database

```bash
docker compose up -d postgres
 
# Wait for PostgreSQL to be healthy
docker compose logs -f postgres
```

{% hint style="info" %}
Watch for the message indicating PostgreSQL is ready to accept connections.
{% endhint %}

3. Start the Pentaho Server.

```bash
docker compose up -d pentaho-server
 
# Monitor startup progress
docker compose logs -f pentaho-server
```

{% hint style="info" %}
The Pentaho Server typically takes 2-3 minutes for first-time initialization. Watch for the message:

```
Server startup in [X] milliseconds
```

{% endhint %}
{% endtab %}
{% endtabs %}

3. Verify container status.

```bash
docker compose ps
```

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
make status
```

4. Run validation script.

```bash
cd
cd ~/Pentaho-Server-PostgreSQL/scripts && ./validate-deployment.sh
```

<figure><img src="/files/S5KL94OfbrSIbxldNKZo" alt=""><figcaption><p>Validate Deployment</p></figcaption></figure>

5. Open a web browser and navigate to:

{% embed url="<http://localhost:8090/pentaho>" %}

5. Login with the default credentials:

| Username | Admin    |
| -------- | -------- |
| Password | password |

6. Enter the Licensing Server URL

<figure><img src="/files/mPIIpegM6yi49gZneyuV" alt=""><figcaption><p>Enter licensing details</p></figcaption></figure>
{% endtab %}

{% tab title="5. Backup & Recovery" %}
{% hint style="info" %}

#### Backup & Recovery

Implement regular backups to protect your Pentaho data and configuration.
{% endhint %}

1. Create a compressed backup of Pentaho databases.

```bash
./scripts/backup-postgres.sh
 
# Backups are saved to backups/ directory with timestamp
# Example: backups/pentaho-postgres-backup-20260113-143022.sql.gz

```

<figure><img src="/files/qAkw6qnIbq7ekO3Qi32E" alt=""><figcaption><p>Backup script</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# K3s

Lightweight Kubernetes ..

{% hint style="info" %}

#### K3s

K3s is a lightweight, certified Kubernetes distribution designed for resource-constrained environments, edge computing, IoT devices, and development scenarios. Created by Rancher Labs (now part of SUSE), it packages everything needed to run Kubernetes into a single binary under 100MB.

This makes K3s significantly lighter than standard Kubernetes while maintaining full compatibility with Kubernetes APIs and features. It requires minimal memory (512MB minimum) and offers simplified operations with reduced dependencies.

K3s comes with built-in components like the Traefik ingress controller, local storage provisioner, and service load balancer. It's perfect for development, CI/CD, edge deployments, and ARM devices, yet remains production-ready with high availability capabilities.
{% endhint %}

<figure><img src="/files/EacVdhTBnAe5SovaYLL4" alt=""><figcaption><p>K3s &#x26; Pentaho Server Architecture</p></figcaption></figure>

{% tabs %}
{% tab title="K3s Core Components" %}
{% hint style="info" %}

#### K3s Core Components

**Control Plane Components**

The control plane includes the API Server (Kubernetes API endpoint for cluster management), Controller Manager (manages core control loops for replication, endpoints, and namespaces), and Scheduler (assigns pods to nodes based on resource availability).

K3s can use either etcd or lightweight SQLite as the datastore for cluster state, making it more flexible than standard Kubernetes.

**Node Components**

Each node runs the Kubelet agent which manages pod lifecycle. Containerd is built-in as the container runtime for running containers.

Kube-proxy manages network proxying and service networking across the cluster.

**Built-in Add-ons**

K3s includes Traefik Ingress Controller for routing external HTTP/HTTPS traffic to services. The Local Path Provisioner enables dynamic persistent volume provisioning using local storage.

CoreDNS provides cluster DNS for service discovery. The Service Load Balancer manages LoadBalancer-type services without requiring external cloud provider integrations.

**Networking**

Flannel serves as the default CNI (Container Network Interface) plugin for pod networking. Network Policies control traffic flow between pods and services for enhanced security.
{% endhint %}

<figure><img src="/files/EVLlvzn29Zwc46a77WSJ" alt=""><figcaption><p>K3s Components</p></figcaption></figure>
{% endtab %}

{% tab title="Pentaho Server Pod" %}
{% hint style="info" %}

#### Pentaho Server Pod

* Runs Tomcat application server with Pentaho Server 11
* Built on Debian Trixie Slim with OpenJDK 21 JRE
* Multi-stage Docker build for optimized image size
* Exposed internally on port 8080
* Includes readiness and liveness probes for health monitoring
* Resource limits configured for CPU and memory stability
  {% endhint %}

<figure><img src="/files/jKxqFpXEO1aau9V5KoO1" alt=""><figcaption><p>Pentaho Pod</p></figcaption></figure>
{% endtab %}

{% tab title="PostgreSQL Pod" %}
{% hint style="info" %}
Provides relational database backend for three critical databases:

* **Jackrabbit** (jcr\_user): Java Content Repository storing all Pentaho content (reports, dashboards, data sources, transformations, jobs)
* **Quartz** (pentaho\_user): Scheduler managing jobs, triggers, calendars, and execution history
* **Hibernate** (hibuser): Security configuration, audit logging, user sessions, plus two specialized schemas:
  * **pentaho\_dilogs**: ETL execution logging with job logs, transformation metrics, and step performance data
  * **pentaho\_operations\_mart**: Dimensional data mart for platform analytics with dimension and fact tables
* Data persisted through PersistentVolumeClaim to survive restarts
* Automated initialization via ConfigMap-mounted SQL scripts
  {% endhint %}

<figure><img src="/files/k3QPX9NXsDew1x4jccTn" alt=""><figcaption><p>PostgreSQL Pod</p></figcaption></figure>

<figure><img src="/files/64dertqWlxEri5oIxnDf" alt=""><figcaption><p><em>PostgreSQL Database Architecture</em></p></figcaption></figure>

Pentaho Server requires three separate databases, each serving a distinct purpose:

<table><thead><tr><th width="128" valign="top">Database</th><th width="133" valign="top">Owner</th><th valign="top">Purpose &#x26; Contents</th></tr></thead><tbody><tr><td valign="top">jackrabbit</td><td valign="top">jcr_user</td><td valign="top">Java Content Repository (JCR) - Stores all Pentaho content including reports, dashboards, data sources, analysis schemas, and user files. This is the primary content storage for the Pentaho repository.</td></tr><tr><td valign="top">quartz</td><td valign="top">pentaho_user</td><td valign="top">Quartz Scheduler - Manages all scheduled jobs, triggers, and calendars. Contains tables for job definitions (QRTZ6_JOB_DETAILS), triggers (QRTZ6_TRIGGERS), execution history, and cluster coordination locks.</td></tr><tr><td valign="top">hibernate</td><td valign="top">hibuser</td><td valign="top">Hibernate Repository - Hosts security configuration, audit logging, user session data, and contains two additional schemas: pentaho_dilogs (ETL execution logging) and pentaho_operations_mart (analytics data mart).</td></tr></tbody></table>

{% hint style="info" %}
The hibernate database contains specialized schemas for operational monitoring:

**pentaho\_dilogs**: Captures detailed ETL execution information including job logs, transformation logs, step performance metrics, and error records. Essential for debugging data integration workflows and monitoring pipeline health.

**pentaho\_operations\_mart**: A dimensional data mart for analytics on Pentaho usage. Contains dimension tables (DIM\_DATE, DIM\_TIME, DIM\_EXECUTOR) and fact tables (FACT\_EXECUTION, FACT\_STEP\_EXECUTION) for analyzing platform utilization, performance trends, and user activity.

For production deployments, implement regular backups of the repository-data Docker volume. The jackrabbit database is the most critical as it contains all user content. Consider using pg\_dump for logical backups or volume snapshots for full recovery options.
{% endhint %}
{% endtab %}

{% tab title="Network" %}
{% hint style="info" %}

#### Netwoking

**Internal Communication:**

* Both pods run within the `pentaho` namespace
* ClusterIP services provide stable internal DNS names
* PostgreSQL accessible at `postgresql.pentaho.svc.cluster.local:5432`
* Pentaho Server accessible at `pentaho-server.pentaho.svc.cluster.local:8080`

**External Access:**

* Traefik Ingress Controller routes external traffic to Pentaho Server
* Configurable hostname and path-based routing
* Optional TLS/SSL termination support
  {% endhint %}

Tomcat manages connection pools defined in context.xml. Each pool serves a specific purpose:

<table><thead><tr><th valign="top">Pool Name</th><th valign="top">Connection Target</th><th valign="top">Used For</th></tr></thead><tbody><tr><td valign="top">jdbc/Hibernate</td><td valign="top">repository:5432/hibernate</td><td valign="top">Security, Users, Roles</td></tr><tr><td valign="top">jdbc/Quartz</td><td valign="top">repository:5432/quartz</td><td valign="top">Job Scheduling</td></tr><tr><td valign="top">jdbc/jackrabbit</td><td valign="top">repository:5432/jackrabbit</td><td valign="top">Content Repository</td></tr><tr><td valign="top">jdbc/Audit</td><td valign="top">repository:5432/hibernate</td><td valign="top">Audit Logging</td></tr><tr><td valign="top">jdbc/live_logging_info</td><td valign="top">repository:5432/hibernate</td><td valign="top">ETL Runtime Logs</td></tr><tr><td valign="top">jdbc/PDI_Operations_Mart</td><td valign="top">repository:5432/hibernate</td><td valign="top">Operations Analytics</td></tr></tbody></table>
{% endtab %}

{% tab title="Storage ConfigMaps" %}
{% hint style="info" %}

#### Storage

**PersistentVolumeClaims (PVCs):**

* `postgres-pvc`: PostgreSQL data directory (`/var/lib/postgresql/data`)
* `pentaho-pvc`: Pentaho solutions and data directories

**Storage Class:**

* Uses K3s's built-in `local-path` storage provisioner
* Provisions volumes on node's local filesystem
* Automatic volume creation and binding

**ConfigMaps:**

* Database initialization scripts (5 SQL files)
* Pentaho configuration settings (JVM parameters, Tomcat settings)

**Secrets:**

* PostgreSQL credentials (postgres\_password, pentaho\_user, pentaho\_password)
* JDBC connection strings
* Base64-encoded for security
  {% endhint %}
  {% endtab %}
  {% endtabs %}

{% hint style="danger" %}
Before you begin the K3s deployment, ensure you have completed the Setup: [Pentaho Containers](/pentaho-11-installation-en/setup/pentaho-containers)
{% endhint %}

Run through the following steps to deploy Pentaho Server on a single-node K3s with PostgreSQL 15 repository.

{% tabs %}
{% tab title="1. Prepare Environment " %}
{% hint style="info" %}

#### Prepare Environment

The "Prepare Environment" section outlines the initial setup steps required before deploying Pentaho Server 11 on K3s:

* copying the deployment assets to your home directory,
* staging the Pentaho Server Enterprise Edition ZIP file, verifying the file is in place,
* confirming K3s is properly installed and running
  {% endhint %}

{% hint style="danger" %}
Ensure you have downloaded: `pentaho-server-ee-11.0.0.0-237.zip`
{% endhint %}

1. Create directory & copy over assets.

```bash
cd
cd ~/Workshop--Installation/Pentaho-Containers/K3s/Pentaho-K3s-PostgreSQL/scripts
chmod +x
./install-to-home.sh
```

2. Copy over the pentaho-server-ee-11.0.0.0-237.zip /docker/stagedArtefacts directory.

{% hint style="info" %}
If you have deployed an Archive Pentaho Server then copy from:

`/opt/pentaho/software/pentaho-server-ee-version`

Otherwise download package from the [Pentaho Customer Portal](https://support.pentaho.com/hc/en-us).
{% endhint %}

```bash
cd
cd ~/Pentaho-K3s-PostgreSQL/docker-build/stagedArtifacts
cp /opt/pentaho/software/server/pentaho-server-ee-11.0.0.0-237.zip . 
```

3. Verify that the file.

```bash
cd
cd ~/Pentaho-K3s-PostgreSQL/docker-build/stagedArtifacts
ls -al
```

4. Verify the K3s installation.

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL/scripts
./verify-k3s.sh
```

```sh
# Normal verification (what you'd run most of the time)
./scripts/verify-k3s.sh

# Detailed output for debugging
./scripts/verify-k3s.sh --verbose

# Silent mode for scripts/automation
./scripts/verify-k3s.sh --quiet

# Get help
./scripts/verify-k3s.sh --help
```

<figure><img src="/files/pC5a2GBcYdAsVU73TqAR" alt=""><figcaption><p>verify-k3s.sh</p></figcaption></figure>

5. Pentaho Server requires a valid license. The `.env` file contains a LICENSE\_URL pointing to the Flexera license server. Ensure your license entitlements are active before deployment.

{% hint style="warning" %}
Without a valid license, Pentaho Server will start but many features will be disabled. Verify your license status before proceeding with production deployments.
{% endhint %}

**Key Differences**

<table><thead><tr><th width="167">Aspect</th><th>Docker Deployment</th><th>K3s Deployment</th></tr></thead><tbody><tr><td><strong>Orchestration</strong></td><td>Docker Compose</td><td>Kubernetes (K3s)</td></tr><tr><td><strong>Configuration</strong></td><td><code>.env</code> file + <code>docker-compose.yml</code></td><td>Kubernetes manifests (YAML)</td></tr><tr><td><strong>Secrets</strong></td><td>Docker secrets or Vault</td><td>Kubernetes Secrets</td></tr><tr><td><strong>Networking</strong></td><td>Docker bridge network</td><td>K3s cluster network + Traefik Ingress</td></tr><tr><td><strong>Storage</strong></td><td>Docker volumes</td><td>PersistentVolumeClaims (PVCs)</td></tr><tr><td><strong>Scaling</strong></td><td>Manual (<code>docker compose up --scale</code>)</td><td>Declarative (<code>replicas</code> in deployment)</td></tr><tr><td><strong>Health Checks</strong></td><td>Docker HEALTHCHECK</td><td>Kubernetes readiness/liveness probes</td></tr><tr><td><strong>Init Scripts</strong></td><td>Volume mount to <code>/docker-entrypoint-initdb.d</code></td><td>ConfigMap mounted to PostgreSQL pod</td></tr></tbody></table>
{% endtab %}

{% tab title="2. Preflight Tasks" %}
{% hint style="info" %}

#### Pre-flight Taks

The Pre-flight Tasks section outlines the essential preparation steps needed before deploying Pentaho Server 11 in K3s containers.

**Configure Environment Variables**

Edit the `.env.example` file within the `docker-build/` directory with your deployment-specific settings. This includes Pentaho version identifier and Docker image name/tag.

Configure PostgreSQL database credentials and connection parameters. Set JVM memory allocation with minimum heap (default 4GB) and maximum heap (default 8GB).

Add your Enterprise Edition license server URL if applicable. Configure build options including image edition (EE/CE), plugin detection, and registry push settings.

Once configured, copy this template to `.env` for use by the build process.

**softwareOverride Directory**

The `softwareOverride/` directory within `docker-build/` provides a mechanism for customizing Pentaho configurations. These customizations get baked into the Docker image during the build process.

Files are organized in numbered directories and processed in alphabetical order.

The **1\_drivers/** directory contains the PostgreSQL JDBC driver (included by default), and you can place additional JDBC drivers here.

The **2\_repository/** directory holds database connection configurations for Jackrabbit (JCR), Quartz (scheduler), and Hibernate repositories. The **3\_security/** directory is empty in this K3s deployment since there's no Vault integration.

The **4\_others/** directory contains modified Tomcat scripts (startup.sh, setenv.sh), server.xml, and other application-level configurations.

You can optionally upgrade the PostgreSQL JDBC driver by downloading from Maven Central or copying from the workshop's database drivers collection. Place the updated driver in `softwareOverride/1_drivers/tomcat/lib/`.
{% endhint %}

***

**Configure .env**

1. Edit the .env.template

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
nano .env.template
```

2. Enter the following details:

<table><thead><tr><th width="261" valign="top">Variable</th><th width="229" valign="top">Default</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top">PENTAHO_VERSION</td><td valign="top">11.0.0.0-237</td><td valign="top">Pentaho Server version</td></tr><tr><td valign="top">EDITION</td><td valign="top">ee</td><td valign="top">Enterprise version</td></tr><tr><td valign="top">INCLUDE_DEMO</td><td valign="top">1</td><td valign="top">Include demo data</td></tr><tr><td valign="top">IMAGE_TAG</td><td valign="top">pentaho/pentaho-server:11.0.0.0-237</td><td valign="top">Docker image tag</td></tr><tr><td valign="top">PENTAHO_MIN_MEMORY</td><td valign="top">4096m</td><td valign="top">JVM minimum heap size</td></tr><tr><td valign="top">PENTAHO_MAX_MEMORY</td><td valign="top">8192m</td><td valign="top">JVM maximum heap size</td></tr><tr><td valign="top">PENTAHO_DI_JAVA_OPTIONS</td><td valign="top">"-Dfile.encoding=utf8 -Djava.awt.headless=true"</td><td valign="top"></td></tr><tr><td valign="top">PENTAHO_IMAGE_NAME</td><td valign="top">pentaho/pentaho-server</td><td valign="top">Docker image name</td></tr><tr><td valign="top">TZ</td><td valign="top">America/NY</td><td valign="top">Time Zone of server</td></tr><tr><td valign="top">DB_TYPE</td><td valign="top">postgres</td><td valign="top"></td></tr><tr><td valign="top">DB_HOST</td><td valign="top">postgres</td><td valign="top"></td></tr><tr><td valign="top">DB_PORT</td><td valign="top">5432</td><td valign="top">PostgreSQL HTTP port</td></tr><tr><td valign="top">PUSH_TO REGISTRY</td><td valign="top">false</td><td valign="top">Pushes direct to K3s Regsitry</td></tr><tr><td valign="top">LOAD_INTO_K3S</td><td valign="top">true</td><td valign="top">Loads directly to K3s</td></tr><tr><td valign="top">RUN TESTS</td><td valign="top">true</td><td valign="top"></td></tr><tr><td valign="top">LICENSE_URL</td><td valign="top">(empty)</td><td valign="top">EE license server URL</td></tr></tbody></table>

3. Save:

```
CTRL + o
Enter
CTRL + x
```

4. Create .env

```bash
cd
cd ~/Pentaho-Server-PostgreSQL
cp .env.template .env
```

***

**softwareOverride**

{% hint style="info" %}
The `softwareOverride/` directory provides a powerful mechanism to customize Pentaho Server without modifying the core installation. Files are copied into the Pentaho installation during container startup, processed in alphabetical order by directory name.
{% endhint %}

````
```
softwareOverride/
├── 1_drivers/           # JDBC drivers and data connectors
│   ├── tomcat/lib/
│   │   └── postgresql-42.x.x.jar    # PostgreSQL JDBC driver (included)
│   └── pentaho-solutions/drivers/    # Big data drivers (.kar files)
├── 2_repository/        # Database repository configuration
│   ├── pentaho-solutions/system/
│   │   ├── hibernate/hibernate-settings.xml
│   │   ├── jackrabbit/repository.xml
│   │   └── scheduler-plugin/quartz/quartz.properties
│   └── tomcat/webapps/pentaho/META-INF/context.xml
├── 3_security/          # Authentication and authorization
│   └── pentaho-solutions/system/
│       ├── applicationContext-spring-security-hibernate.properties
│       └── applicationContext-spring-security-memory.xml
├── 4_others/            # Tomcat, defaults, and miscellaneous
│   ├── pentaho-solutions/system/
│   │   ├── defaultUser.spring.properties
│   │   ├── pentaho.xml
│   │   └── security.properties
│   └── tomcat/
│       ├── bin/startup.sh
│       └── webapps/pentaho/WEB-INF/web.xml
└── 99_exchange/         # User data exchange (not auto-processed)
```
````

The PostgreSQL JDBC driver is included in the Pentaho distribution. If you need to upgrade:

1. Download from [Maven Central](https://repo1.maven.org/maven2/org/postgresql/postgresql/)
2. Place in `softwareOverride/1_drivers/tomcat/lib/`

Or

Copy from Workshop--Installation/'Database Drivers'/

```bash
cd
cd ~/Workshop--Installation/'Database Drivers'
cp postgresql-42.7.8.jar ~/Pentaho-Server-PostgreSQL/softwareOverride/1_drivers/tomcat/lib
```

{% endtab %}

{% tab title="3. Build & Push Image" %}
{% hint style="info" %}

#### Build & Push Pentaho Image

The build.sh script is an automated build wrapper that:

* Validates prerequisites - Checks Docker is installed
* Checks for required files - Verifies Pentaho ZIP exists in stagedArtifacts/
* Detects plugins automatically - Finds PAZ, PIR, PDD plugins
* Confirms build - Shows what will be built and asks for confirmation
* Runs docker build - Executes the build with proper arguments
* Shows image info - Displays image size and details after build
* Optional: Tests image - Runs basic container test (you can skip this)
* Optional: Pushes to registry - Pushes to Docker registry (only with -p flag)

This is the **recommended approach** for building Pentaho Docker images. It uses a single `.env` file to configure everything - similar to the Docker Compose deployment.
{% endhint %}

You can modify the build with the following options:

<table><thead><tr><th width="144">Option (short)</th><th width="183">Option (Long)</th><th>Description</th></tr></thead><tbody><tr><td>-v</td><td>--version VERSION</td><td>Pentaho version (default: 11.0.0.0-237)</td></tr><tr><td>-t</td><td>--tag TAG</td><td>Docker image tag (default: pentaho/pentaho-server:VERSION)</td></tr><tr><td>-e</td><td>--edition EDITION</td><td>ee or ce (default: ee)</td></tr><tr><td>-d</td><td>--demo</td><td>Include demo content (default: no)</td></tr><tr><td>-p</td><td>--push</td><td>Push to registry after build</td></tr><tr><td>-h</td><td>--help</td><td>Push to registry after build</td></tr></tbody></table>

1. Build & Push the Pentaho Server Image directly into K3s Registry.

```bash
# Build with defaults based on .env
cd
cd ~/Pentaho-K3s-PostgreSQL/docker-build
./build.sh
```

<figure><img src="/files/nptbkMr3YorCtJ8IZHeq" alt=""><figcaption><p>Build &#x26; Push Pentaho Server image</p></figcaption></figure>

{% hint style="info" %}
**1. Docker Image Build:** The deployment uses a custom-built Pentaho Server container image:

```
docker-build/
├── Dockerfile (multi-stage build)
├── build.sh (automated build script)
├── stagedArtifacts/ (Pentaho Server ZIP)
└── softwareOverride/ (configuration overlays)
    ├── 1_drivers/ (PostgreSQL JDBC driver)
    ├── 2_repository/ (database configs)
    ├── 3_security/ (authentication)
    └── 4_others/ (Tomcat customizations)
```

The `build.sh` script:

* Validates prerequisites and required files
* Detects plugins automatically (PAZ, PIR, PDD)
* Executes multi-stage Docker build
* Optionally pushes to K3s image store

**2. Configuration Management:**

* `.env` file contains deployment-specific settings (versions, credentials, JVM memory)
* `softwareOverride/` directory provides configuration overlays processed in numbered order
* PostgreSQL JDBC driver included by default, with option to upgrade
* Custom Tomcat scripts for container startup optimization
  {% endhint %}
  {% endtab %}

{% tab title="4. Helm Charts" %}
{% hint style="info" %}

#### Helm Charts

Helm is the package manager for Kubernetes, often referred to as "apt/yum for Kubernetes." It simplifies the deployment and management of Kubernetes applications by:

**Packaging**: Bundling related Kubernetes resources together

**Templating**: Parameterizing manifests for reusability across environments

**Versioning**: Managing application versions and upgrades

**Release Management**: Tracking deployments and enabling rollbacks
{% endhint %}

<figure><img src="/files/0DjVb3LBK5WFIHeLpDon" alt=""><figcaption><p>Helm Charts</p></figcaption></figure>

{% tabs %}
{% tab title="1. Directories" %}
{% hint style="info" %}

#### Directory Layout

Below is an explanation of the pentaho directory responsible for a Helm deployment.
{% endhint %}

```
pentaho/                          # Chart root directory
├── Chart.yaml                    # Chart metadata (name, version, description)
├── values.yaml                   # Default configuration values
├── templates/                    # Kubernetes manifest templates
│   ├── _helpers.tpl              # Template helper functions (not rendered)
│   ├── NOTES.txt                 # Post-installation instructions
│   ├── namespace.yaml            # Namespace creation
│   ├── secret.yaml               # Sensitive data (passwords, keys)
│   ├── configmap-*.yaml          # Configuration data
│   ├── pvc.yaml                  # Persistent volume claims
│   ├── *-deployment.yaml         # Pod deployments
│   ├── *-service.yaml            # Service definitions
│   └── ingress.yaml              # Ingress routing rules
└── files/                        # Non-template files (SQL scripts, configs)
    └── db_init/                  # PostgreSQL initialization scripts
```

**Chart.yaml**

{% hint style="info" %}
This file defines the chart's identity, version, and metadata used by Helm.

It serves as the "package definition" for the Helm chart, similar to package.json (npm) or Chart.lock (Helm dependencies).

Role: Chart.yaml tells Helm what this chart is, what version it is, what application it deploys, and provides searchable metadata for chart repositories.

Helm uses this file to track chart versions, manage dependencies, and display information when users search for or install charts.
{% endhint %}

**values.yaml**

{% hint style="info" %}
In a K3s (lightweight Kubernetes) deployment, the `values.yaml` file serves as the central configuration file for Helm charts. It defines default parameters and settings that customize how an application or service is deployed to the cluster - things like replica counts, image versions, resource limits, service types, ingress rules, environment variables, and persistent storage configurations.

When you run `helm install` or `helm upgrade`, Helm merges the values from this file with the chart's templates to generate the final Kubernetes manifests. You can override specific values at deploy time using `--set` flags or by supplying a custom values file with `-f`, making it a flexible mechanism for managing environment-specific configurations (e.g., dev vs. production) without modifying the underlying chart templates.
{% endhint %}

***

**templates**

<table><thead><tr><th width="201">YAML</th><th>Description</th></tr></thead><tbody><tr><td>namespace.yaml</td><td>Creates an isolated Kubernetes namespace to contain all Pentaho resources</td></tr><tr><td>secret.yaml</td><td>Stores sensitive database credentials (passwords) encrypted in Kubernetes</td></tr><tr><td>config-*.yaml</td><td>Configures Pentaho environment variables (JVM memory, database settings, paths, timezone) Contains SQL scripts to initialize PostgreSQL databases (jackrabbit, quartz, hibernate) on first startup</td></tr><tr><td>pvc.yaml</td><td>Requests persistent storage volumes for PostgreSQL data and optional Pentaho data/solutions</td></tr><tr><td>*-deployment.yaml</td><td>Deploys Pentaho Business Analytics Server with init container, health probes, and resource limits. Deploys PostgreSQL 15 database server with automatic initialization and persistent storage</td></tr><tr><td>*-service.yaml</td><td>Routes external HTTP/HTTPS traffic to Pentaho Server via Traefik ingress controller. Exposes PostgreSQL port 5432 as a stable DNS endpoint for Pentaho to connect to.</td></tr><tr><td>ingress.yaml</td><td>Routes external HTTP/HTTPS traffic to Pentaho Server via Traefik ingress controller</td></tr></tbody></table>
{% endtab %}

{% tab title="2. Deploy" %}
{% hint style="info" %}

#### Deploy

The architecture consists of two main pods running within a dedicated Pentaho namespace:

a Pentaho Server pod (Tomcat on Debian with OpenJDK 21) and

a PostgreSQL pod hosting three essential databases - Jackrabbit for content storage, Quartz for job scheduling, and Hibernate for security and audit logging.
{% endhint %}

{% hint style="danger" %}
Before you proceed, ensure you have completed Steps 1 - 3. You should have a Pentaho Server image + PostgreSQL 15 repository images pushed to the K3s repository.
{% endhint %}

1. Quick check.

```bash
# Check Helm version (requires 3.0+)
helm version

# Check Kubernetes cluster
kubectl cluster-info
kubectl get nodes

# Check available storage classes
kubectl get storageclass
```

2. Check The Pentaho Image is available.

```bash
# Verify
sudo k3s ctr images ls | grep pentaho
```

<figure><img src="/files/XPFgpS5Z0eHs6zU3jxtl" alt=""><figcaption><p>Pentaho Server image in K3s repository</p></figcaption></figure>

***

{% hint style="info" %}

#### Default Deployment

The deployment workflow progresses through four stages:

* preparing the environment by staging the Pentaho Enterprise Edition package and verifying K3s.
* configuring environment variables and software overrides during pre-flight tasks.
* building and pushing a custom Docker image into the K3s container runtime.
* finally deploying the full stack using either Helm charts or an automated `deploy.sh` script that orchestrates namespace creation, secrets, ConfigMaps, persistent storage, and Traefik ingress routing.

Also covers the Helm chart structure, a suite of helper scripts for backup, health checks, resource monitoring, and deployment validation, along with multiple access methods including port-forwarding, hostname-based ingress, and direct node IP.
{% endhint %}

1. Install using Helm charts.

```bash
cd 
cd ~/Pentaho-K3s-PostgreSQL/helm-chart
helm install pentaho ./pentaho
```

<figure><img src="/files/iYBvQxNujSaRCwuzYFnY" alt=""><figcaption><p>Deploy Pentaho Server - Helm Chart</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Helper Scripts" %}
{% hint style="info" %}

#### Helper Scripts

The K3s Pentaho deployment includes a collection of helper scripts in the `scripts/` directory for day-to-day operations and maintenance:

* **backup-postgres.sh:** Creates timestamped, compressed dumps of all Pentaho databases (Jackrabbit, Quartz, Hibernate) using `pg_dump` executed inside the PostgreSQL pod.
* **restore-postgres.sh:** Restores Pentaho databases from backup files, useful for disaster recovery, environment cloning, or migrating data between clusters.
* **health-check.sh:** Performs a quick runtime health check covering pod readiness, service availability, database connectivity, and a live HTTP test against the Pentaho login page.
* **validate-deployment.sh:** Runs a comprehensive audit across six categories: namespace, pods, services, PersistentVolumeClaims, ConfigMaps, and ingress configuration.
* **monitor-resources.sh:** Tracks CPU and memory usage across all Pentaho pods using `kubectl top` to help identify resource constraints or optimization opportunities.
* **monitor-postgres.sh:** Monitors PostgreSQL-specific metrics including connection counts, active queries, table sizes, and overall database health.
* **verify-k3s.sh:** Validates the underlying K3s infrastructure (node status, storage classes, Traefik ingress, and core components) before attempting a Pentaho deployment.
* **Makefile:** Provides convenience commands such as `make health`, `make status`, `make logs`, `make port-forward`, and `make destroy` for streamlined cluster management.
  {% endhint %}

{% tabs %}
{% tab title="1. Directories" %}
{% hint style="info" %}

#### Directory Layout

This K3s deployment configuration provides several important capabilities:

* Completely self-contained Kubernetes deployment on lightweight K3s
* Automated database initialization with PostgreSQL SQL scripts
* Kubernetes-native health checks and startup ordering
* Persistent volume claims for database and Pentaho content
* Docker image build process with multi-stage optimization
* Resource limits (CPU/memory) for stability
* Production-ready Kubernetes manifest templates
* PostgreSQL JDBC driver included
* Easy backup and restore procedures via utility scripts
* Ingress configuration for Traefik routing
  {% endhint %}

***

**Root Directory Files**

```
Pentaho-K3s-PostgreSQL/
├── deploy.sh                    # Main deployment script
├── destroy.sh                   # Cleanup script
├── Makefile                     # Quick commands (make help)
├── README.md                    # This file
├── DEPLOYMENT.md                # Detailed deployment guide
├── K3s-INSTALLATION.md          # K3s setup instructions
```

{% hint style="info" %}
**Documentation Files:**

**README.md** - The main entry point documentation providing project overview, quick start instructions, prerequisites, and general usage information for the K3s deployment workshop.

**DEPLOYMENT.md** - Detailed deployment guide covering the complete K3s deployment workflow, including pre-requisites, step-by-step instructions, and post-deployment verification procedures.

**K3s-INSTALLATION.md** - K3s setup and installation instructions covering system preparation, K3s installation, networking configuration, and storage class setup required before deploying Pentaho.

**Orchestration & Deployment:**

**deploy.sh** - Automated deployment script that orchestrates the complete K3s deployment workflow including namespace creation, secret generation, manifest application, and service readiness checks.

**destroy.sh** - Cleanup script that safely removes all K3s resources including deployments, services, persistent volume claims, and namespaces, useful for redeployment or teardown scenarios.

**Makefile** - Contains convenience command targets for common K3s operations like deploying, destroying, checking status, and viewing logs. Users can run `make help` to see all available commands.
{% endhint %}

***

**docker-build**

```
├── docker-build/                # Docker image build
│   ├── build.sh                 # Build script
│   ├── Dockerfile               # Image definition
│   ├── README.md                # Complete build documentation
│   ├── QUICK-START.md           # Quick build guide
│   ├── ENV-CONFIGURATION.md     # .env configuration guide
│   ├── .env.example             # Configuration template
│   ├── test-compose.yml         # Local Docker testing
│   ├── entrypoint/              # Container startup scripts
│   │   ├── docker-entrypoint.sh
│   │   └── start-pentaho-docker.sh
│   ├── softwareOverride/        # Config overlays (baked into image)
│   │   ├── 1_drivers/           # PostgreSQL JDBC driver
│   │   ├── 2_repository/        # Database configs
│   │   │   └── README.md
│   │   ├── 3_security/          # (empty - no Vault)
│   │   └── 4_others/            # Modified Tomcat scripts
│   │       └── README.md
│   └── stagedArtifacts/         # Place Pentaho ZIP here
│       └── README.md
```

{% hint style="info" %}
The **docker-build/** directory contains all components needed to build the Pentaho Server container image that will be deployed to K3s:

**Documentation Files:**

**README.md** - Complete build documentation covering the Docker image build process, multi-stage build architecture, configuration options, troubleshooting, and best practices for building Pentaho Server images for K3s deployment.

**QUICK-START.md** - Quick build guide providing step-by-step instructions for users who want to quickly build and test the Pentaho image without reading the complete documentation. Includes common build commands and typical workflows.

**ENV-CONFIGURATION.md** - Comprehensive configuration reference guide for the `.env.example` file, detailing all available environment variables, their purposes, default values, and how they affect the Docker build process and resulting image.

**Build & Configuration Files:**

**build.sh** - Automated build wrapper script that validates prerequisites, checks for required files (Pentaho ZIP in `stagedArtifacts/`), detects plugins automatically (PAZ, PIR, PDD), confirms the build with the user, executes `docker build`, shows image info after build, and optionally pushes to a registry with the `-p` flag.

**.env.example** - Configuration template file containing all available environment variables for the Docker build process. Users copy this to `.env` and customize values for their specific deployment needs including Pentaho version, image tags, and build options.

**Dockerfile** - Multi-stage build definition using `debian:trixie-slim` as the base image with OpenJDK 21 JRE. Creates optimized images by separating the build environment from the runtime environment, reducing final image size while maintaining all necessary Pentaho components. The multi-stage approach minimizes security vulnerabilities and improves build efficiency.

**test-compose.yml** - Local Docker Compose testing environment that allows you to test the built Docker image locally before deploying to K3s. This is useful for validating configuration changes, testing custom plugins, or debugging startup issues without the overhead of a full K3s deployment.

**Container Startup Scripts:**

**entrypoint/** - Directory containing container initialization and startup scripts:

* **docker-entrypoint.sh** - Primary container startup script that executes when the container starts. Handles environment variable processing, configuration file customization from `softwareOverride/`, database connection validation, health checks, and orchestrates the Pentaho Server startup sequence.
* **start-pentaho-docker.sh** - Pentaho-specific startup script that manages Tomcat initialization, JVM configuration, memory settings, and starts the Pentaho Server services. This script is called by `docker-entrypoint.sh` after environment preparation is complete.

**Configuration Overlays:**

**softwareOverride/** - Configuration overlays directory that gets baked into the Docker image during the build process. Files are organized in numbered directories and processed in alphabetical order to ensure proper application sequence:

* **1\_drivers/** - PostgreSQL JDBC driver (included by default) for database connectivity. Additional JDBC drivers or data connectors can be placed here.
* **2\_repository/** - Database connection configurations for all Pentaho repositories including Jackrabbit (JCR), Quartz (scheduler), and Hibernate (security/audit). Contains a **README.md** explaining the repository configuration files and their purposes.
* **3\_security/** - Empty in this K3s deployment since HashiCorp Vault integration is not used. In production environments, this would contain authentication, authorization, and security configuration files.
* **4\_others/** - Modified Tomcat scripts (startup.sh, setenv.sh), server.xml, web.xml, and other application-level configurations. Contains a **README.md** documenting the custom Tomcat modifications and their purposes.

**Staged Artifacts:**

**stagedArtifacts/** - Staging directory where users place the Pentaho Server installation package (`pentaho-server-ee-11.0.0.0-237.zip`) before building the Docker image. Contains a **README.md** with instructions on where to obtain the Pentaho package and how to stage it properly.
{% endhint %}

***

**db\_init\_postgres**

```
├── db_init_postgres/                       # PostgreSQL init SQL scripts
│   ├── 1_create_jcr_postgresql.sql
│   ├── 2_create_quartz_postgresql.sql
│   ├── 3_create_repository_postgresql.sql
│   ├── 4_pentaho_logging_postgresql.sql
│   └── 5_pentaho_mart_postgresql.sql
```

{% hint style="info" %}
The **db\_init\_postgres/** directory contains PostgreSQL initialization scripts that create all required Pentaho repository databases. These scripts are mounted into the PostgreSQL container and execute automatically on first startup:

**1\_create\_jcr\_postgresql.sql** - Creates the **Jackrabbit Content Repository (JCR)** database. The JCR stores all Pentaho content including reports, dashboards, data sources, analysis schemas, transformations, jobs, and user files. This is the primary content management system for the Pentaho repository.

**2\_create\_quartz\_postgresql.sql** - Sets up the **Quartz Scheduler** database. Quartz manages all scheduled jobs, triggers, and calendars within Pentaho Server, including report generation schedules, ETL job executions, and other automated processes. Contains critical tables like `QRTZ6_JOB_DETAILS`, `QRTZ6_TRIGGERS`, and execution history.

**3\_create\_repository\_postgresql.sql** - Creates the **Hibernate Repository** database. This stores user authentication data, authorization information, roles, permissions, and other security-related information managed by Pentaho's security subsystem.

**4\_pentaho\_logging\_postgresql.sql** - Establishes the **pentaho\_dilogs** schema within the Hibernate database for audit and Data Integration (DI) logging. Captures detailed ETL execution information including job logs, transformation logs, step performance metrics, and error records. Essential for debugging data integration workflows and monitoring pipeline health.

**5\_pentaho\_mart\_postgresql.sql** - Creates the **pentaho\_operations\_mart** schema within the Hibernate database. This dimensional data mart stores operational analytics about Pentaho Server usage, including dimension tables (`DIM_DATE`, `DIM_TIME`, `DIM_EXECUTOR`) and fact tables (`FACT_EXECUTION`, `FACT_STEP_EXECUTION`) for analyzing platform utilization, performance trends, and user activity patterns.
{% endhint %}

***

**manifests**

```
├── manifests/                          # Kubernetes manifests
│   ├── namespace.yaml                  # Pentaho namespace
│   ├── configmaps/
│   │   ├── pentaho-config.yaml
│   │   └── postgres-init-scripts.yaml
│   ├── pentaho/
│   │   ├── deployment.yaml
│   │   └── service.yaml
│   ├── postgres/
│   │   ├── deployment.yaml
│   │   └── service.yaml
│   ├── secrets/
│   │   └── secrets.yaml                # (gitignored)
│   ├── storage/
│   │   └── pvc.yaml
│   └── ingress/
│       └── ingress.yaml
```

{% hint style="info" %}
The **manifests/** directory contains all Kubernetes resource definitions organized by functional area. These YAML files define the declarative state of your K3s deployment:

**namespace.yaml** - Creates the dedicated `pentaho` namespace to isolate all Pentaho-related resources from other K3s workloads, providing logical separation and resource organization.

**configmaps/** - ConfigMap resources for non-sensitive configuration data:

* **pentaho-config.yaml** - Pentaho Server configuration settings like JVM parameters, Tomcat settings, and application properties
* **postgres-init-scripts.yaml** - ConfigMap containing the five PostgreSQL initialization scripts from `db_init_postgres/` directory, mounted into the PostgreSQL pod

**pentaho/** - Pentaho Server Kubernetes resources:

* **deployment.yaml** - Defines the Pentaho Server deployment including container specifications, resource requests/limits, environment variables, volume mounts, readiness/liveness probes, and replica count
* **service.yaml** - ClusterIP service exposing Pentaho Server on port 8080 within the cluster, providing stable internal DNS and load balancing

**postgres/** - PostgreSQL database Kubernetes resources:

* **deployment.yaml** - Defines the PostgreSQL 15 deployment with container specifications, persistent volume claims for data storage, initialization script mounting, and database configuration
* **service.yaml** - ClusterIP service exposing PostgreSQL on port 5432 within the cluster for Pentaho Server database connections

**secrets/** - Sensitive credential storage:

* **secrets.yaml** - Kubernetes Secret resource containing base64-encoded credentials for PostgreSQL (`postgres_password`, `pentaho_user`, `pentaho_password`) and JDBC connection strings. **This file is gitignored for security.**

**storage/** - Persistent storage definitions:

* **pvc.yaml** - PersistentVolumeClaim definitions for both PostgreSQL data (`postgres-pvc`) and Pentaho solutions/data (`pentaho-pvc`), using K3s's local-path storage class for persistent data across pod restarts

**ingress/** - External access configuration:

* **ingress.yaml** - Traefik Ingress resource defining external HTTP/HTTPS routing rules to expose Pentaho Server outside the K3s cluster, including hostname, path routing, and TLS configuration if applicable
  {% endhint %}

***

**scripts**

```
└── scripts/                     # Utility scripts
    ├── backup-postgres.sh       # Database backup
    ├── restore-postgres.sh      # Database restore
    ├── health-check.sh          # Health check
    ├── monitor-resources.sh     # Resource monitoring
    ├── monitor-postgres.sh      # PostgreSQL monitoring
    ├── validate-deployment.sh   # Deployment validation
    └── verify-k3s.sh            # K3s verification
```

{% hint style="info" %}
The **scripts/** directory contains operational and maintenance utilities for managing the K3s Pentaho deployment:

**Database Management:**

**backup-postgres.sh** - Automated PostgreSQL backup utility that creates compressed dumps of all Pentaho databases (jackrabbit, quartz, hibernate) using `kubectl exec` to run `pg_dump` inside the PostgreSQL pod. Backups are timestamped and compressed with gzip for efficient storage.

**restore-postgres.sh** - Database restoration utility to recover Pentaho databases from backup files. Useful for disaster recovery, environment cloning, or migrating data between K3s clusters. Handles decompression and restoration via `kubectl exec` and `psql`.

**Monitoring & Validation:**

**health-check.sh** - Health check script that verifies both PostgreSQL and Pentaho Server are running and responding correctly. Checks pod status, readiness probes, and performs basic connectivity tests.

**monitor-resources.sh** - Resource monitoring utility that tracks CPU, memory, and storage usage across all Pentaho pods using `kubectl top` and resource metrics, helping identify resource constraints or optimization opportunities.

**monitor-postgres.sh** - PostgreSQL-specific monitoring script that checks database connection counts, active queries, table sizes, and database health metrics via SQL queries executed in the PostgreSQL pod.

**validate-deployment.sh** - Comprehensive deployment validation script that confirms all K3s resources are properly created, pods are running, services are accessible, persistent volumes are bound, and the entire deployment is operational.

**verify-k3s.sh** - K3s infrastructure verification script that checks K3s installation, node status, storage classes, Traefik ingress controller, and core K3s components before attempting Pentaho deployment.
{% endhint %}

***

**Key Differences: K3s vs Docker**

<table><thead><tr><th width="167">Aspect</th><th width="262">Docker Deployment</th><th>K3s Deployment</th></tr></thead><tbody><tr><td><strong>Orchestration</strong></td><td>Docker Compose</td><td>Kubernetes (K3s)</td></tr><tr><td><strong>Configuration</strong></td><td><code>.env</code> file + <code>docker-compose.yml</code></td><td>Kubernetes manifests (YAML)</td></tr><tr><td><strong>Secrets</strong></td><td>Docker secrets or Vault</td><td>Kubernetes Secrets</td></tr><tr><td><strong>Networking</strong></td><td>Docker bridge network</td><td>K3s cluster network + Traefik Ingress</td></tr><tr><td><strong>Storage</strong></td><td>Docker volumes</td><td>PersistentVolumeClaims (PVCs)</td></tr><tr><td><strong>Scaling</strong></td><td>Manual (<code>docker compose up --scale</code>)</td><td>Declarative (<code>replicas</code> in deployment)</td></tr><tr><td><strong>Health Checks</strong></td><td>Docker HEALTHCHECK</td><td>Kubernetes readiness/liveness probes</td></tr><tr><td><strong>Init Scripts</strong></td><td>Volume mount to <code>/docker-entrypoint-initdb.d</code></td><td>ConfigMap mounted to PostgreSQL pod</td></tr></tbody></table>
{% endtab %}

{% tab title="2. Deploy" %}
{% hint style="info" %}

#### Deployment execution

The `deploy.sh` script automates the entire workflow:

* Verifies K3s is running
* Creates namespace
* Applies all manifests in correct order
* Monitors pod startup
* Validates service readiness
* Provides deployment summary with access URLs

**1. Namespace Creation:**

```bash
kubectl create namespace pentaho
```

Creates isolated logical environment for all Pentaho resources.

**2. Secret Generation:**

```bash
kubectl apply -f manifests/secrets/secrets.yaml
```

Stores PostgreSQL credentials and JDBC connection strings as Kubernetes Secrets.

**3. Storage Provisioning:**

```bash
kubectl apply -f manifests/storage/pvc.yaml
```

Creates PersistentVolumeClaims for PostgreSQL data and Pentaho content.

**4. ConfigMap Creation:**

```bash
kubectl apply -f manifests/configmaps/
```

Mounts PostgreSQL initialization scripts and Pentaho configuration.

**5. PostgreSQL Deployment:**

```bash
kubectl apply -f manifests/postgres/
```

Deploys PostgreSQL pod with:

* Mounted init scripts (automatic database creation on first startup)
* Persistent volume for data
* Health checks and resource limits
* ClusterIP service for internal connectivity

**6. Pentaho Server Deployment:**

```bash
kubectl apply -f manifests/pentaho/
```

Deploys Pentaho pod with:

* Custom Docker image
* Environment variables from ConfigMap and Secrets
* Volume mounts for solutions/data
* Readiness/liveness probes
* ClusterIP service

**7. Ingress Configuration:**

```bash
kubectl apply -f manifests/ingress/ingress.yaml
```

Configures Traefik routing for external access.
{% endhint %}

1. Run the deploy.sh

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL
./deploy.sh
```

<figure><img src="/files/kXsF6LfIqA4fFdsjSV5x" alt=""><figcaption><p>Deploy Pentaho Server</p></figcaption></figure>

{% hint style="info" %}
This unified script handles the complete deployment process for Pentaho Server on K3s, including:

* Docker image import into K3s container runtime
* Kubernetes resource creation (namespace, secrets, configmaps, storage)
* PostgreSQL database deployment
* Pentaho Server deployment
* Ingress configuration
* Health checks and status reporting

```
Usage:
./deploy.sh # Fresh deployment with image import
./deploy.sh --skip-import # Deploy without importing image
./deploy.sh --update-only # Only update existing deployment
./deploy.sh --clean # Remove old images before deploying
```

Prerequisites:

* K3s installed and running
* Docker image built: pentaho/pentaho-server:11.0.0.0-237
* kubectl configured to access K3s cluster
* sudo access for K3s containerd operations
  {% endhint %}

***

**Quick Commands with Makefile**

```bash
make help           # Show all available commands
make full-deploy    # Build, import, and deploy (complete workflow)
make health         # Run health check
make status         # Show deployment status
make logs           # View Pentaho logs
make port-forward   # Access Pentaho at localhost:8080
make destroy        # Remove deployment
```

***

There's also bunch of scripts that will help validate the deployment:

{% tabs %}
{% tab title="Validate Deployment" %}
{% hint style="info" %}
Comprehensive post-deployment validation script that verifies all components of the Pentaho K3s deployment are properly configured and running correctly. What It Checks (6 Categories)

**Namespace**

* Verifies pentaho namespace exists

**Pods**

* PostgreSQL pod is Running
* Pentaho Server pod is Running
* Shows current status if not running

**Services**

* PostgreSQL service exists
* Pentaho Server service exists

**PersistentVolumeClaims (3 PVCs)**

* postgres-data-pvc (10Gi) - Database files
* pentaho-data-pvc (10Gi) - Pentaho
* data pentaho-solutions-pvc (5Gi) - Solutions repository
* All must be in "Bound" status

**ConfigMaps**

* pentaho-config - Environment variables
* postgres-init - Database initialization scripts

**Ingress**

* pentaho-ingress - Traefik routing configuration

**Database Connectivity Tests**

* Connects to PostgreSQL pod
* Tests all 3 Pentaho databases:

\* jackrabbit - JCR content repository

\* quartz - Job scheduler

\* hibernate - Configuration repository

* Runs SELECT 1 query on each
  {% endhint %}

1. Run the following `validate-deployment.sh` script.

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL/scripts
./validate-deployment.sh
```

<figure><img src="/files/h92PfcOx0oftsJ91r5oA" alt=""><figcaption><p>Validate deployment</p></figcaption></figure>
{% endtab %}

{% tab title="Health Check" %}
{% hint style="info" %}

#### Health Check

Quick health check script for running Pentaho deployment - faster and lighter than full validation, focused on runtime health status.

Namespace

* Verifies pentaho namespace exists
* Exits immediately if namespace missing (critical)

Pod Readiness

* PostgreSQL pod is ready (not just running)
* Pentaho Server pod is ready
* Checks `containerStatuses[0].ready status`

Services

* PostgreSQL service exists
* Pentaho Server service exists

Database Connectivity

* PostgreSQL is responding to queries

All 3 databases exist:

\* jackrabbit

\* quartz

\* hibernate

* Uses `psql -lqt` to list databases

Web Application Health

* Live HTTP test to Pentaho login page
* Uses port-forward to access service
* Expects HTTP 200 response
* Tests: <http://localhost:8080/pentaho/Login>

Resource Usage

* Shows CPU/memory usage via `kubectl top pods`
* Gracefully handles missing metrics-server
  {% endhint %}

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL/scripts
./health-check.sh
```

<figure><img src="/files/A4q13oCh3GQGYub5OQ7O" alt=""><figcaption><p>Health Check</p></figcaption></figure>
{% endtab %}

{% tab title="Monitor Resources" %}
{% hint style="info" %}

#### Monitor Resources

x
{% endhint %}

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL/scripts
./validate-deployment.sh
```

<figure><img src="/files/jruslK6W2rgeqInZsOeX" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Pentaho Server" %}
{% hint style="info" %}

#### Accessing Pentaho Server

{% endhint %}

**Port Forward (Recommended for Testing/Development)**

The simplest method using kubectl to forward a local port to the Pentaho service:

```bash
kubectl port-forward -n pentaho svc/pentaho-server 8080:8080
```

**Access URL:** `http://localhost:8080/pentaho`

You can also use an alternate port if 8080 is busy:

```bash
kubectl port-forward -n pentaho svc/pentaho-server 9080:8080
# Then access: http://localhost:9080/pentaho
```

***

**Ingress with Hostname (pentaho.local)**

Uses K3s's built-in Traefik ingress controller with DNS-style access:

**Setup:**

```bash
echo "10.0.0.1 pentaho.local" | sudo tee -a /etc/hosts
```

*(Replace 10.0.0.1 with your actual node IP)*

**Access URL:** `http://pentaho.local/pentaho`

***

**Ingress via Direct Node IP (No DNS Required)**

Access directly through any cluster node's IP address without configuring /etc/hosts:

**Access URL:** `http://<node-ip>/pentaho`

This works because the ingress includes a path-based rule that doesn't require a hostname.

***

**Makefile Convenience Command**

The project includes a Makefile shortcut:

```bash
make port-forward
```

This automatically sets up port forwarding to localhost:8080.

***

**Default Credentials**

| Username | Password |
| -------- | -------- |
| admin    | password |

⚠️ Change these for production deployments!

***

**Quick Reference**

| Method              | Best For            | URL                             |
| ------------------- | ------------------- | ------------------------------- |
| Port Forward        | Development/Testing | `http://localhost:8080/pentaho` |
| Ingress (hostname)  | Production with DNS | `http://pentaho.local/pentaho`  |
| Ingress (direct IP) | Testing without DNS | `http://<node-ip>/pentaho`      |
| {% endtab %}        |                     |                                 |
| {% endtabs %}       |                     |                                 |
| {% endtab %}        |                     |                                 |
| {% endtabs %}       |                     |                                 |
| {% endtab %}        |                     |                                 |
| {% endtabs %}       |                     |                                 |


# Ubuntu Pentaho Lab

Setup Pentaho Server + Plugins on Ubuntu ..

{% hint style="info" %}

#### **Pentaho Lab**

Pentaho Data Integration is a client-based tool commonly installed and configured to run on Windows 11.

There are several licensing options, for these workshops we will be installing a Enterprise Edition. This will give you the opportunity to try out building a complete solution - automated data pipelines + analytics ..
{% endhint %}

<figure><img src="/files/3VBLH4tZldXqRmYd2Lwo" alt=""><figcaption><p>Pentaho Tiers</p></figcaption></figure>

{% hint style="danger" %}
The following steps are intended for setting up a Pentaho Lab environment and need to be completed in order to complete the Workshops.

Ensure you have downloaded the Workshop--Installation:

```bash
cd
git clone https://github.com/jporeilly/Workshop--Installation
```

To install git:

```bash
sudo apt install git
```

{% endhint %}

{% hint style="info" %}
**Prerequisites**

* Ubuntu 24.04 LTS system (physical or virtual machine)
* User account with sudo privileges
* Internet connection
* Basic familiarity with Linux command line
  {% endhint %}

{% tabs %}
{% tab title="Docker" %}
{% hint style="info" %}

#### Docker

Docker is a platform that enables developers to package applications and their dependencies into lightweight, portable containers. Containers ensure that applications run consistently across different computing environments, from development laptops to production servers. This workshop will guide you through the complete process of installing Docker Engine on Ubuntu 24.04 LTS (Noble Numbat).
{% endhint %}

1. Before installing Docker, update your existing package list.

```bash
sudo apt update && sudo apt upgrade
```

2. Install packages that allow apt to use repositories over HTTPS.

```bash
sudo apt install -y ca-certificates curl gnupg lsb-release
```

3. Create a directory for keyrings and add Docker's GPG key.

```bash
sudo install -m 0755 -d /etc/apt/keyrings
```

```bash
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
```

```bash
sudo chmod a+r /etc/apt/keyrings/docker.gpg
```

4. Add the Docker repository to your apt sources.

```bash
echo \
  "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu \
  $(. /etc/os-release && echo "$VERSION_CODENAME") stable" | \
  sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
```

5. Now that the Docker repository is added, update the package index.

```bash
sudo apt update && sudo apt upgrade
```

6. Install Docker Engine, containerd, and Docker Compose.

```bash
sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
```

7. Check that Docker is installed correctly by checking the version.

```bash
docker --version
```

You should see output similar to (Nov 2025):

```
Docker version 29.0.2, build 810xxxx
```

8. Verify that Docker Engine is running.

```bash
sudo systemctl status docker
```

The service should show as "active (running)".

9. Quit.

```bash
q
```

10. Test your Docker installation by running the hello-world container.

```bash
sudo docker run hello-world
```

{% hint style="info" %}
This command downloads a test image and runs it in a container. If successful, you'll see a message confirming that Docker is working correctly.
{% endhint %}

***

{% hint style="info" %}

#### Without Sudo

By default, Docker requires sudo privileges. To run Docker commands without sudo.
{% endhint %}

1. Add your user to the docker group.

```bash
sudo usermod -aG docker $USER
```

2. Apply the new group membership (or log out and back in).

```bash
newgrp docker
```

3. Verify you can run Docker without sudo.

```bash
docker run hello-world
```

4. Ensure Docker starts automatically when the system boots.

```bash
sudo systemctl enable docker.service
sudo systemctl enable containerd.service
```

***

{% hint style="info" %}

#### Verification & Testing

To confirm everything is working properly, run the following commands:
{% endhint %}

Check Docker version:

```bash
docker version
```

View Docker system information:

```bash
docker info
```

List running containers:

```bash
docker ps
```

List all containers (including stopped ones):

```bash
docker ps -a
```

List downloaded images:

```bash
docker images
```

***

{% hint style="info" %}

#### Common Commands

Here are essential Docker commands you'll use regularly:

* `docker pull <image>` - Download an image from Docker Hub
* `docker images` - List all local images
* `docker run <image>` - Create and start a container from an image
* `docker ps` - List running containers
* `docker ps -a` - List all containers
* `docker stop <container>` - Stop a running container
* `docker rm <container>` - Remove a stopped container
* `docker rmi <image>` - Remove an image
* `docker logs <container>` - View container logs
* `docker exec -it <container> bash` - Access a running container's shell
  {% endhint %}
  {% endtab %}

{% tab title="MySQL Container" %}
{% hint style="info" %}

#### **Docker Compose - MySQL**

The pentaho\_admin user only has READ permission for the Steel Wheels - sampledata database. The administrator account has been removed.

As you'll be running through CRUID database operations we need to deploy a sampledata database - Docker container, granting all privileges to an admin user.
{% endhint %}

{% file src="/files/ZPzBcxeeelTSjxlWcXXK" %}

{% file src="/files/OUwzKjbSmSC7ujrABdVT" %}

{% file src="/files/mpgAHyKvJ8nEtyuoGlGw" %}

1. Run the following script to create a MySQL folder and copy the required files.

```bash
cd
cd ~/Workshop--Installation/MySQL
# Make it executable
chmod +x copy_mysql.sh
./copy_mysql.sh
```

2. Check the Directory has been created and the files copied over.
3. Execute the docker-compose script to create the container.

```bash
cd
cd ~/MySQL
# Make it executable
chmod +x run_mysql_compose.sh
./run_mysql_compose.sh
```

<figure><img src="/files/U9jdl5M0fzncYYwOd43D" alt="" width="500"><figcaption><p>Deploy MySQL Docker container</p></figcaption></figure>

4. Check the container is up and running in Docker.

```
docker ps
```

<figure><img src="/files/1aHF1xzgyctmxIzj8xBn" alt=""><figcaption><p>Docker MySQL containers</p></figcaption></figure>
{% endtab %}

{% tab title="sampledata" %}
{% hint style="info" %}

#### **Sampledata**

Next on the list is to create the sampledata database.
{% endhint %}

<figure><img src="/files/crBrZ7OsNMJjcWi3WDZ1" alt=""><figcaption><p>Relationship Diagram</p></figcaption></figure>

{% tabs %}
{% tab title="1. sampledata\_schema.sql" %}
{% hint style="info" %}

#### **sampledata\_schema.sql**

This script creates a comprehensive relational database structure for a sample business application. It's designed to model a sales and order management system for a company that sells various products.
{% endhint %}

{% hint style="info" %}

#### **Database Setup**

* Creates a database named

  ```
  sampledata
  ```

  with UTF-8 character set
* Sets up users with appropriate permissions
* Configures SQL mode for better data integrity
  {% endhint %}

{% hint style="info" %}

#### **Tables**

**OFFICES**: Stores company office locations with address details

**EMPLOYEES**: Contains employee information with relationships to offices and reporting structure

**CUSTOMERS**: Stores customer information including contact details and credit limits

**PRODUCTS**: Contains product catalog with inventory and pricing information

**ORDERS**: Tracks customer orders with status and dates

**ORDERDETAILS**: Contains line items for each order with quantity and price

**PAYMENTS**: Records customer payments with amounts and dates

**ORDERFACT**: A fact table for order analytics

**CUSTOMER\_W\_TER**: Extended customer information with territory

**DIM\_TIME**: Time dimension table for reporting

**DEPARTMENT\_MANAGERS**: Stores department manager information

**QUADRANT\_ACTUALS**: Contains budget vs. actual financial data with a generated VARIANCE column

**TRIAL\_BALANCE**: Financial accounting data
{% endhint %}

{% hint style="info" %}

#### **Views**

**customer\_order\_summary**: Summarizes orders and spending by customer

**product\_performance**: Analyzes product sales metrics including revenue and profit

**employee\_sales\_performance**: Tracks sales performance by employee

**monthly\_sales\_trend**: Shows sales trends over time by month

**product\_inventory\_status**: Categorizes products by inventory levels

**customer\_payment\_history**: Summarizes customer payment activity and balances
{% endhint %}

{% hint style="info" %}

#### **Stored Procedures**

**GetCustomerOrders**: Retrieves orders for a specific customer

**UpdateProductStock**: Updates product inventory levels

**GetProductSalesByQuarter**: Analyzes quarterly product sales

**GetTopCustomersByRegion**: Identifies top customers by region

**GetInventoryValueByProductLine**: Calculates inventory metrics by product line

**Triggers**

**before\_order\_insert**: Validates date constraints on orders

**before\_payment\_insert**: Ensures payment amounts are positive
{% endhint %}

{% hint style="info" %}

#### **Events**

* **daily\_maintenance**: Scheduled task for database maintenance
  {% endhint %}

1. Execute the following command to create the schema.

```bash
cd 
cd ~/MySQL
cat sampledata_schema.sql | docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

{% hint style="info" %}
This command is importing SQL schema data into a MySQL database running in a Docker container. Here's a breakdown:

This command reads the SQL file:

```powershell
cat sampledata_schema.sql
```

Pipes (forwards) the file contents to the next command:

```
|
```

This executes a command in a running Docker container:

```
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

{% endhint %}

2. You can check the sampledata database & tables with the following commands.

Show databases:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" -e "SHOW DATABASES;"
```

Show tables:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SHOW TABLES;"
```

Show table columns:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "DESCRIBE table_name;"
```

{% endtab %}

{% tab title="2. sampledata\_data.sql" %}
{% hint style="info" %}

#### **sampledata\_data.sql**

This script populates the database with sample data to demonstrate the functionality of the schema.
{% endhint %}

{% hint style="info" %}

#### **Reference Data**

* Office locations across different regions
* Employee hierarchy with job titles
* Product catalog organized by product lines
  {% endhint %}

{% hint style="info" %}

#### **Transactional Data**

* Customer records with contact information
* Order history with dates and status
* Order details with quantities and prices
* Payment records
  {% endhint %}

{% hint style="info" %}

#### **Financial Data**

* Budget vs. actual figures in QUADRANT\_ACTUALS
* Trial balance accounting data
  {% endhint %}

{% hint style="info" %}

#### **Data Characteristics**

* Realistic business scenarios with varied order statuses
* Comprehensive product catalog with descriptions and pricing
* Hierarchical employee structure with reporting relationships
* Time-based data spanning multiple years for trend analysis
* Financial data suitable for budgeting and variance analysis
  {% endhint %}

{% hint style="info" %}

#### **Notable Features**

* Data follows referential integrity constraints
* Proper handling of NULL values where appropriate
* Realistic pricing and quantity values
* Generated columns (like VARIANCE) are excluded from direct inserts
* Orders are sequenced to satisfy foreign key constraints
  {% endhint %}

1. Execute the following command to load the data into the sampledata tables.

```bash
cd 
cd ~/MySQL
cat sampledata_data.sql | docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

2. You can use the following commands to check that the data has loaded.

To count the number of rows in a specific table:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT COUNT(*) FROM table_name;"
```

To view the first few rows from a table:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT * FROM table_name LIMIT 10;"
```

To check counts for all tables:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT TABLE_NAME, TABLE_ROWS FROM INFORMATION_SCHEMA.TABLES WHERE TABLE_SCHEMA = 'sampledata';"
```

To get a summary of tables and their statuses:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SHOW TABLE STATUS;"
```

<figure><img src="/files/pAXjaVcDvoMVfC7kBVCd" alt=""><figcaption><p>Check Tables</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="DBeaver" %}
{% hint style="info" %}

#### DBeaver

Your going to need a database management tool. DBeaver Community is a free, open-source database management tool for personal projects.
{% endhint %}

1. Simpliest option is to download & install from Snapstore.

```bash
cd
sudo apt update && sudo apt upgrade
sudo apt install snapd
sudo snap install dbeaver-ce
```

Or

Go to the official [DBeaver download page](https://dbeaver.io/download/)

{% embed url="<https://dbeaver.io/files/dbeaver-ce_latest_amd64.deb>" %}

Or

To install that DEB file.

```bash
wget https://dbeaver.io/debs/dbeaver.gpg.key -O /tmp/dbeaver.gpg.key
sudo gpg --dearmor -o /usr/share/keyrings/dbeaver.gpg /tmp/dbeaver.gpg.key
echo "deb [signed-by=/usr/share/keyrings/dbeaver.gpg] https://dbeaver.io/debs/dbeaver-ce /" | sudo tee /etc/apt/sources.list.d/dbeaver.list
sudo apt update
sudo apt install dbeaver-ce
```

4. Pin DBeaver to Dash - bottom toolbar.

***

{% hint style="info" %}

#### **MySQL Database**

If you have completed the previous 3 requirements, then you should have a MySQL Docker container, exposed on port:3306 with sampledata databse.
{% endhint %}

1. Launch DBeaver and Select: MySQL.

<figure><img src="/files/f04HxAgnFLcJJfT7g1Kc" alt=""><figcaption><p>MySQL</p></figcaption></figure>

2. Configure the connection with the following properties:

Username: root or pentaho\_user

Password: password

<figure><img src="/files/YlrpU0j9USa8MQwptL9F" alt=""><figcaption><p>Configure &#x26; Test MySQL connection - sampledata</p></figcaption></figure>

{% hint style="warning" %}
You may need to download the supported version of the database driver.

Also enable: allowPublicKeyRetrieval
{% endhint %}

<figure><img src="/files/3nu7BbPVJSa69szd9xkW" alt=""><figcaption><p>Enable: allowPublicKeyRetrieval</p></figcaption></figure>

3. Test the connection.

<figure><img src="/files/zi3B824EcsskmFjDo7Am" alt=""><figcaption><p>Test connection</p></figcaption></figure>

4. Expand: databases > sampledata > Tables

<figure><img src="/files/IpfBGicoVihojG7VvsCo" alt=""><figcaption><p>Customer Data</p></figcaption></figure>

5. Open a SQL window and run a test query.

```sql
select * from CUSTOMERS
where COUNTRY = 'USA' and CITY = 'NYC';
```

<figure><img src="/files/5B25msWIcQHRygRLGOTv" alt=""><figcaption><p>SQL query - NYC Customers</p></figcaption></figure>
{% endtab %}
{% endtabs %}

<details>

<summary>General Troubleshooting (click to expand)</summary>

**Issue: "permission denied" errors**

* Solution: Ensure your user is in the docker group and you've logged out/in or run `newgrp docker`

**Issue: Docker service won't start**

* Solution: Check logs with `sudo journalctl -u docker.service`

**Issue: Cannot connect to Docker daemon**

* Solution: Ensure Docker service is running with `sudo systemctl start docker`

</details>


# Windows Pentaho Lab

{% hint style="info" %}

#### **Pentaho Lab**

Pentaho Data Integration is a client-based tool commonly installed and configured to run on Windows 11.

There are several licensing options, for these workshops we will be installing a Enterprise Edition. This will give you the opportunity to try out building a complete solution - automated data pipelines + analytics ..
{% endhint %}

<figure><img src="/files/3VBLH4tZldXqRmYd2Lwo" alt=""><figcaption><p>Pentaho Tiers</p></figcaption></figure>

{% hint style="danger" %}
The following steps are intended for setting up a Pentaho Lab environment and need to be completed in order to complete the Workshops.

Ensure you have downloaded the Workshop--Installation

```
cd \
git clone https://github.com/jporeilly/Workshop--Installation
```

{% endhint %}

{% tabs %}
{% tab title="Docker Desktop" %}
{% hint style="info" %}

#### Docker Desktop

Docker Desktop is an application for Windows, macOS, and Linux that provides an easy-to-use interface for developing and running containerized applications. It bundles the Docker Engine, Docker CLI, Docker Compose, Kubernetes, and other essential tools into a single package with a graphical user interface.

Docker Desktop simplifies container management by handling the underlying virtualization automatically, allowing developers to build, test, and deploy applications in isolated, portable containers without worrying about environment configuration differences. It's particularly popular among developers who want to ensure their applications run consistently across different environments, from local development machines to production servers.
{% endhint %}

{% embed url="<https://www.docker.com/products/docker-desktop/>" %}

1. Download the Docker Desktop installer.

{% embed url="<https://desktop.docker.com/win/main/amd64/Docker%20Desktop%20Installer.exe?utm_campaign=dd-smartbutton&utm_location=module&utm_medium=webreferral&utm_source=docker>" %}
Link to download Docker Desktop
{% endembed %}

2. Navigate to: Downloads
3. Double-click: `Docker Desktop Installer.exe` to run the installer.

By default, Docker Desktop is installed at `C:\Program Files\Docker\Docker`.

{% hint style="danger" %}
When prompted, ensure the **Use WSL 2 instead of Hyper-V** option on the Configuration page is selected.

On systems that support only one backend, Docker Desktop automatically selects the available option.
{% endhint %}

<figure><img src="/files/BJEAzzcnABlV9PclrX2k" alt=""><figcaption><p>Use WSL 2</p></figcaption></figure>

3. Close to complete the installation process.

***

{% hint style="info" %}

#### **Docker User**

If your administrator account is different to your user account, you must add the user to the docker-users group:
{% endhint %}

1. Run Computer Management as an administrator.
2. Navigate to **Local Users and Groups** > **Groups** > **docker-users**.
3. Right-click to add the user to the group.

<figure><img src="/files/1OQMqQy99e6XelX5T02n" alt=""><figcaption><p>Add User to docker group</p></figcaption></figure>

4. Sign out and sign back in for the changes to take effect.
   {% endtab %}

{% tab title="MySQL Container" %}
{% hint style="info" %}

#### **Docker Compose - MySQL**

The pentaho\_admin user only has READ permission for the Steel Wheels - sampledata database. The administrator account has been removed.

As you'll be running through CRUID database operations we need to deploy a sampledata database - Docker container, granting all privileges to an admin user.
{% endhint %}

{% file src="/files/bjYKAUpo2Iqv3hbTaoZm" %}

{% file src="/files/LJeMuPGqwic6ImlsjRa5" %}

{% file src="/files/mpgAHyKvJ8nEtyuoGlGw" %}

1. Run the following script to create a MySQL folder and copy the required files.

```ps1
cd \
cd C:\Workshop--Installation\MySQL
.\copy-mysql.ps1
```

2. Check the Directory has been created and the files copied over.
3. Execute the docker-compose script to create the container.

```powershell
cd \
cd C:\MySQL
.\run-docker-mysql.ps1
```

<figure><img src="/files/qiyHtYcWq3Z2bjeyKhyU" alt=""><figcaption><p>Deploy MySQL</p></figcaption></figure>

4. Check the container is up and running in Docker Desktop.

<figure><img src="/files/93fiySPB6INJWPr1aGzC" alt=""><figcaption><p>mysql docker containers</p></figcaption></figure>
{% endtab %}

{% tab title="sampledata" %}
{% hint style="info" %}

#### **Sampledata**

Next on the list is to create the sampledata database.
{% endhint %}

<figure><img src="/files/crBrZ7OsNMJjcWi3WDZ1" alt=""><figcaption><p>Relationship Diagram</p></figcaption></figure>

{% tabs %}
{% tab title="1. sampledata\_schema.sql" %}
{% hint style="info" %}

#### **sampledata\_schema.sql**

This script creates a comprehensive relational database structure for a sample business application. It's designed to model a sales and order management system for a company that sells various products.
{% endhint %}

{% hint style="info" %}

#### **Database Setup**

* Creates a database named

  ```
  sampledata
  ```

  with UTF-8 character set
* Sets up users with appropriate permissions
* Configures SQL mode for better data integrity
  {% endhint %}

{% hint style="info" %}

#### **Tables**

**OFFICES**: Stores company office locations with address details

**EMPLOYEES**: Contains employee information with relationships to offices and reporting structure

**CUSTOMERS**: Stores customer information including contact details and credit limits

**PRODUCTS**: Contains product catalog with inventory and pricing information

**ORDERS**: Tracks customer orders with status and dates

**ORDERDETAILS**: Contains line items for each order with quantity and price

**PAYMENTS**: Records customer payments with amounts and dates

**ORDERFACT**: A fact table for order analytics

**CUSTOMER\_W\_TER**: Extended customer information with territory

**DIM\_TIME**: Time dimension table for reporting

**DEPARTMENT\_MANAGERS**: Stores department manager information

**QUADRANT\_ACTUALS**: Contains budget vs. actual financial data with a generated VARIANCE column

**TRIAL\_BALANCE**: Financial accounting data
{% endhint %}

{% hint style="info" %}

#### **Views**

**customer\_order\_summary**: Summarizes orders and spending by customer

**product\_performance**: Analyzes product sales metrics including revenue and profit

**employee\_sales\_performance**: Tracks sales performance by employee

**monthly\_sales\_trend**: Shows sales trends over time by month

**product\_inventory\_status**: Categorizes products by inventory levels

**customer\_payment\_history**: Summarizes customer payment activity and balances
{% endhint %}

{% hint style="info" %}

#### **Stored Procedures**

**GetCustomerOrders**: Retrieves orders for a specific customer

**UpdateProductStock**: Updates product inventory levels

**GetProductSalesByQuarter**: Analyzes quarterly product sales

**GetTopCustomersByRegion**: Identifies top customers by region

**GetInventoryValueByProductLine**: Calculates inventory metrics by product line

**Triggers**

**before\_order\_insert**: Validates date constraints on orders

**before\_payment\_insert**: Ensures payment amounts are positive
{% endhint %}

{% hint style="info" %}

#### **Events**

* **daily\_maintenance**: Scheduled task for database maintenance
  {% endhint %}

1. Execute the following command to create the schema.

```powershell
cd \
cd MySQL
Get-Content sampledata_schema.sql | docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

{% hint style="info" %}
This command is importing SQL schema data into a MySQL database running in a Docker container. Here's a breakdown:

This command reads the SQL file:

```powershell
Get-Content sampledata_schema.sql
```

Pipes (forwards) the file contents to the next command:

```
|
```

This executes a command in a running Docker container:

```
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

{% endhint %}

2. You can check the sampledata database & tables with the following commands.

Show databases:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" -e "SHOW DATABASES;"
```

Show tables:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SHOW TABLES;"
```

Show table columns:

```powershell
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "DESCRIBE table_name;"
```

<figure><img src="/files/qHh8I5vHVQP5NtT9YrGA" alt=""><figcaption><p>Check database &#x26; tables</p></figcaption></figure>
{% endtab %}

{% tab title="2. sampledata\_data.sql" %}
{% hint style="info" %}

#### **sampledata\_data.sql**

This script populates the database with sample data to demonstrate the functionality of the schema.
{% endhint %}

{% hint style="info" %}

#### **Reference Data**

* Office locations across different regions
* Employee hierarchy with job titles
* Product catalog organized by product lines
  {% endhint %}

{% hint style="info" %}

#### **Transactional Data**

* Customer records with contact information
* Order history with dates and status
* Order details with quantities and prices
* Payment records
  {% endhint %}

{% hint style="info" %}

#### **Financial Data**

* Budget vs. actual figures in QUADRANT\_ACTUALS
* Trial balance accounting data
  {% endhint %}

{% hint style="info" %}

#### **Data Characteristics**

* Realistic business scenarios with varied order statuses
* Comprehensive product catalog with descriptions and pricing
* Hierarchical employee structure with reporting relationships
* Time-based data spanning multiple years for trend analysis
* Financial data suitable for budgeting and variance analysis
  {% endhint %}

{% hint style="info" %}

#### **Notable Features**

* Data follows referential integrity constraints
* Proper handling of NULL values where appropriate
* Realistic pricing and quantity values
* Generated columns (like VARIANCE) are excluded from direct inserts
* Orders are sequenced to satisfy foreign key constraints
  {% endhint %}

1. Execute the following command to load the data into the sampledata tables.

```powershell
cd \
cd MySQL
Get-Content sampledata_data.sql | docker exec -i mysql-database-1 mysql -u root -p"password" sampledata
```

2. You can use the following commands to check that the data has loaded.

To count the number of rows in a specific table:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT COUNT(*) FROM table_name;"
```

To view the first few rows from a table:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT * FROM table_name LIMIT 10;"
```

To check counts for all tables:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SELECT TABLE_NAME, TABLE_ROWS FROM INFORMATION_SCHEMA.TABLES WHERE TABLE_SCHEMA = 'sampledata';"
```

To get a summary of tables and their statuses:

```docker
docker exec -i mysql-database-1 mysql -u root -p"password" sampledata -e "SHOW TABLE STATUS;"
```

<figure><img src="/files/DrXc52UvRyyIgeEBKahh" alt=""><figcaption><p>sampledata table info</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="DBeaver" %}
{% hint style="info" %}

#### DBeaver

Your going to need a database management tool. DBeaver Community is a free, open-source database management tool for personal projects.
{% endhint %}

1. Go to the official [DBeaver download page](https://dbeaver.io/download/)

{% embed url="<https://dbeaver.io/files/dbeaver-ce-latest-x86_64-setup.exe>" %}

2. Navigate to Downloads & double-click on: `dbeaver-ce-25.2.5-x86_64-setup.exe`
3. Follow the installation instructions.
4. Follow the on-screen instructions, clicking "Next" and agreeing to the license agreement to proceed.
5. Choose your desired installation options (e.g., for all users or the current user).

<figure><img src="/files/dEM3Hp5QMtAbnUTL6qlB" alt=""><figcaption></figcaption></figure>

6. Complete the installation process.

***

{% hint style="info" %}

#### **MySQL Database**

If you have completed the previous 3 requirements, then you should have a MySQL Docker container, exposed on port:3306 with sampledata databse.
{% endhint %}

1. Launch DBeaver and Select: MySQL.

<figure><img src="/files/f04HxAgnFLcJJfT7g1Kc" alt=""><figcaption><p>MySQL</p></figcaption></figure>

2. Configure the connection with the following properties:

Username: root or pentaho\_user

Password: password

<figure><img src="/files/YlrpU0j9USa8MQwptL9F" alt=""><figcaption><p>Configure &#x26; Test MySQL connection - sampledata</p></figcaption></figure>

{% hint style="warning" %}
You may need to download the supported version of the database driver.

Also enable: allowPublicKeyRetrieval
{% endhint %}

<figure><img src="/files/3nu7BbPVJSa69szd9xkW" alt=""><figcaption><p>Enable: allowPublicKeyRetrieval</p></figcaption></figure>

3. Test the connection.

<figure><img src="/files/zi3B824EcsskmFjDo7Am" alt=""><figcaption><p>Test connection</p></figcaption></figure>

4. Expand: databases > sampledata > Tables

<figure><img src="/files/IpfBGicoVihojG7VvsCo" alt=""><figcaption><p>Customer Data</p></figcaption></figure>

5. Open a SQL window and run a test query.

```sql
select * from CUSTOMERS
where COUNTRY = 'USA' and CITY = 'NYC';
```

<figure><img src="/files/5B25msWIcQHRygRLGOTv" alt=""><figcaption><p>SQL query - NYC Customers</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Pentaho Containers

{% hint style="info" %}

#### Pentaho Containers

Pentaho Server offers flexible deployment options to suit a range of infrastructure strategies. For **on-premises** environments, organizations can deploy using Docker Compose for straightforward single-host setups, or orchestrate across clusters using Kubernetes (K8s & K3s) for greater scalability and resilience.

For **cloud (hyperscaler)** deployments, Pentaho provides optimized, prebuilt Docker images purpose-built for Amazon Web Services (EKS/ECR), Microsoft Azure (AKS), and Google Cloud Platform (GKE), allowing teams to leverage managed Kubernetes services and cloud-native storage like Amazon S3. As of Pentaho 11, these images feature standardized installation paths, environment variables, and enhanced entrypoint scripts that support runtime configuration overrides - meaning licenses and configuration files can be injected at startup without rebuilding the image.

A **hybrid** approach is also fully supported, where organizations might run the Pentaho Server on-premises for sensitive workloads while deploying Carte server containers or PDI worker nodes in the cloud to handle burst processing, blending the control of local infrastructure with the elasticity of cloud resources.
{% endhint %}

{% hint style="danger" %}
The following steps are intended for setting up a Pentaho Lab environment and need to be completed in order to complete the Workshops.

Ensure you have downloaded the Workshop--Installation:

```bash
cd
git clone https://github.com/jporeilly/Workshop--Installation
```

To install git:

```bash
sudo apt install git
```

{% endhint %}

Select your container host:

{% tabs %}
{% tab title="Docker" %}
{% hint style="info" %}

#### Docker

Docker is a platform that enables developers to package applications and their dependencies into lightweight, portable containers. Containers ensure that applications run consistently across different computing environments, from development laptops to production servers. This workshop will guide you through the complete process of installing Docker Engine on Ubuntu 24.04 LTS (Noble Numbat).
{% endhint %}

1. Before installing Docker, update your existing package list.

```bash
sudo apt update && sudo apt upgrade
```

2. Install packages that allow apt to use repositories over HTTPS.

```bash
sudo apt install -y ca-certificates curl gnupg lsb-release
```

3. Create a directory for keyrings and add Docker's GPG key.

```bash
sudo install -m 0755 -d /etc/apt/keyrings
```

```bash
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
```

```bash
sudo chmod a+r /etc/apt/keyrings/docker.gpg
```

4. Add the Docker repository to your apt sources.

```bash
echo \
  "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu \
  $(. /etc/os-release && echo "$VERSION_CODENAME") stable" | \
  sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
```

5. Now that the Docker repository is added, update the package index.

```bash
sudo apt update && sudo apt upgrade
```

6. Install Docker Engine, containerd, and Docker Compose.

```bash
sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
```

7. Check that Docker is installed correctly by checking the version.

```bash
docker --version
```

You should see output similar to (Nov 2025):

```
Docker version 29.0.2, build 810xxxx
```

8. Verify that Docker Engine is running.

```bash
sudo systemctl status docker
```

The service should show as "active (running)".

9. Quit.

```bash
q
```

10. Test your Docker installation by running the hello-world container.

```bash
sudo docker run hello-world
```

{% hint style="info" %}
This command downloads a test image and runs it in a container. If successful, you'll see a message confirming that Docker is working correctly.
{% endhint %}

***

{% hint style="info" %}

#### Without Sudo

By default, Docker requires sudo privileges. To run Docker commands without sudo.
{% endhint %}

1. Add your user to the docker group.

```bash
sudo usermod -aG docker $USER
```

2. Apply the new group membership (or log out and back in).

```bash
newgrp docker
```

3. Verify you can run Docker without sudo.

```bash
docker run hello-world
```

4. Ensure Docker starts automatically when the system boots.

```bash
sudo systemctl enable docker.service
sudo systemctl enable containerd.service
```

***

{% hint style="info" %}

#### Verification & Testing

To confirm everything is working properly, run the following commands:
{% endhint %}

Check Docker version:

```bash
docker version
```

View Docker system information:

```bash
docker info
```

List running containers:

```bash
docker ps
```

List all containers (including stopped ones):

```bash
docker ps -a
```

List downloaded images:

```bash
docker images
```

***

{% hint style="info" %}

#### Common Commands

Here are essential Docker commands you'll use regularly:

* `docker pull <image>` - Download an image from Docker Hub
* `docker images` - List all local images
* `docker run <image>` - Create and start a container from an image
* `docker ps` - List running containers
* `docker ps -a` - List all containers
* `docker stop <container>` - Stop a running container
* `docker rm <container>` - Remove a stopped container
* `docker rmi <image>` - Remove an image
* `docker logs <container>` - View container logs
* `docker exec -it <container> bash` - Access a running container's shell
  {% endhint %}
  {% endtab %}

{% tab title="K3s" %}
{% hint style="info" %}

#### K3s

K3s is a lightweight, fully compliant Kubernetes distribution designed for resource-constrained and edge computing environments. It's packaged as a single binary or minimal container image, making it significantly easier to deploy and manage than standard Kubernetes. The distribution uses SQLite3 as its default lightweight datastore, though it also supports etcd3, MySQL, and Postgres for users needing more robust options.

The platform simplifies Kubernetes operations by wrapping all control plane components into a single binary and process. This unified approach automates complex tasks like certificate distribution and TLS configuration, while maintaining security by default with sensible settings for lightweight environments. K3s has minimal external dependencies, requiring only a modern Linux kernel and cgroup mounts to run.

K3s comes "batteries-included" with essential components pre-packaged, eliminating the need for separate installation and configuration. This includes containerd for container runtime, Flannel for networking, CoreDNS for service discovery, Traefik for ingress, and several other critical controllers for load balancing, network policies, storage, and image management. This comprehensive bundle makes K3s ideal for quick cluster creation in edge locations, IoT deployments, CI/CD pipelines, and development environments.
{% endhint %}

{% embed url="<https://docs.k3s.io/>" %}

1. Before installing Docker, update your existing package list.

```bash
sudo apt update && sudo apt upgrade
```

2. Install packages that allow apt to use repositories over HTTPS.

```bash
sudo apt install -y ca-certificates curl gnupg lsb-release
```

3. Disable swap (recommended for Kubernetes).

```bash
# Disable swap immediately
sudo swapoff -a

# Disable swap permanently
sudo sed -i '/ swap / s/^/#/' /etc/fstab

# Verify swap is disabled
free -h | grep Swap
```

4. Enable IP Forwarding.

```bash
sudo tee /etc/sysctl.d/k3s.conf <<EOF 
net.ipv4.ip_forward = 1 
net.bridge.bridge-nf-call-iptables = 1 
net.bridge.bridge-nf-call-ip6tables = 1 
EOF 
sudo sysctl --system
```

5. Configure Firewall.

```bash
# Check if UFW is active
sudo ufw status

# If UFW is active, allow required ports
sudo ufw allow 6443/tcp    # Kubernetes API
sudo ufw allow 8472/udp    # Flannel VXLAN
sudo ufw allow 10250/tcp   # Kubelet
sudo ufw allow 80/tcp      # HTTP
sudo ufw allow 443/tcp     # HTTPS
```

***

Select your deployment options:

{% tabs %}
{% tab title="Single-Node" %}
{% hint style="info" %}

#### **Single-Node Installation**

{% endhint %}

1. Check System requirements.

```bash
# Check system requirements
echo "=== System Information ==="
echo "OS: $(lsb_release -d | cut -f2)"
echo "Kernel: $(uname -r)"
echo "CPU Cores: $(nproc)"
echo "Total RAM: $(free -h | awk '/^Mem:/ {print $2}')"
echo "Available Disk: $(df -h / | awk 'NR==2 {print $4}')"
```

2. Install K3s.

```bash
# Install K3s server (includes agent)
curl -sfL https://get.k3s.io | sh -
```

{% hint style="info" %}
This command:

\- Downloads and installs K3s

\- Starts the K3s service

\- Installs kubectl

\- Configures kubeconfig

```
What gets installed:
- K3s server (control plane + worker)
- Containerd (container runtime)
- Flannel (CNI network plugin)
- CoreDNS (cluster DNS)
- Traefik (ingress controller)
- Local-path provisioner (storage)
- Metrics server
```

{% endhint %}

2. Configure kubectl access.

```bash
# Create kube config directory
mkdir -p ~/.kube

# Copy K3s config to standard location
sudo cp /etc/rancher/k3s/k3s.yaml ~/.kube/config

# Fix permissions
sudo chown $(id -u):$(id -g) ~/.kube/config
chmod 600 ~/.kube/config

# Set KUBECONFIG environment variable
echo 'export KUBECONFIG=~/.kube/config' >> ~/.bashrc
source ~/.bashrc

# Test kubectl access
kubectl version --short
```

4\. Verify Installation

```bash
# Check node status
kubectl get nodes

# Expected output:
# NAME         STATUS   ROLES                  AGE   VERSION
NAME      STATUS   ROLES           AGE     VERSION
pentaho   Ready    control-plane   5d20h   v1.34.3+k3s1

# Check system pods
kubectl get pods -A
```

<figure><img src="/files/pZoz9KsdpdRZeamapURi" alt=""><figcaption><p>Verify K3s</p></figcaption></figure>

5. Verify the K3s installation.

```sh
cd
cd ~/Pentaho-K3s-PostgreSQL
./verify-k3s.sh
```

{% endtab %}

{% tab title="Multi-Node" %}
{% hint style="info" %}

#### Multi-Node Cluster

{% endhint %}

#### Server Node (Control Plane)

```bash
# Install K3s server
curl -sfL https://get.k3s.io | sh -

# Get the node token (needed for agents)
sudo cat /var/lib/rancher/k3s/server/node-token
```

#### Agent Nodes (Workers)

On each worker node:

```bash
# Set variables
K3S_URL="https://<server-ip>:6443"
K3S_TOKEN="<node-token-from-server>"

# Install K3s agent
curl -sfL https://get.k3s.io | K3S_URL=$K3S_URL K3S_TOKEN=$K3S_TOKEN sh -
```

#### Verify Multi-Node Cluster

On the server node:

```bash
kubectl get nodes

# Expected output:
# NAME         STATUS   ROLES                  AGE   VERSION
# server-01    Ready    control-plane,master   5m    v1.29.0+k3s1
# worker-01    Ready    <none>                 2m    v1.29.0+k3s1
# worker-02    Ready    <none>                 1m    v1.29.0+k3s1
```

{% endtab %}

{% tab title="Kubectl Aliases" %}

1. Add helpful aliases.

```bash
# Add helpful aliases
cat >> ~/.bashrc << 'EOF'

# Kubernetes aliases
alias k='kubectl'
alias kgp='kubectl get pods'
alias kgs='kubectl get svc'
alias kgn='kubectl get nodes'
alias kga='kubectl get all'
alias kd='kubectl describe'
alias kl='kubectl logs'
alias kaf='kubectl apply -f'
alias kdf='kubectl delete -f'

# Enable kubectl autocompletion
source <(kubectl completion bash)
complete -F __start_kubectl k
EOF

source ~/.bashrc
```

{% endtab %}

{% tab title="Post-Installation" %}
{% hint style="info" %}

### Post-Installation Configuration

{% endhint %}

1. Verify Default Storage Class

```bash
# Verify storage class
kubectl get storageclass
```

```
# Expected output:
NAME                   PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ALLOWVOLUMEEXPANSION   AGE
local-path (default)   rancher.io/local-path   Delete          WaitForFirstConsumer   false                  5d21h
```

2. Configure Traefik Ingress

```bash
kubectl get pods -n kube-system -l app.kubernetes.io/name=traefik
```

3. Install Helm (Optional but Recommended)

```bash
# Install Helm
curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash

# Verify
helm version
```

4\. Install k9s (Optional - Terminal UI)

```bash
# Install k9s for easier cluster management
curl -sS https://webinstall.dev/k9s | bash

# Run k9s
k9s
```

<figure><img src="/files/j22p8Hlg2y9Sui36Maph" alt=""><figcaption><p>K9s</p></figcaption></figure>

5\. Verify K3s Installation.

```bash
cd
cd ~/Pentaho-K3s-PostgreSQL

# Make script executable
chmod +x scripts/verify-k3s.sh

# Run verification
./scripts/verify-k3s.sh
```

<figure><img src="/files/JSwyChKd1yQ0E2LTTmaP" alt=""><figcaption><p>Verify K3s</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Helm" %}
{% hint style="info" %}

#### Helm Charts

Helm is the package manager for Kubernetes, often referred to as "apt/yum for Kubernetes." It simplifies the deployment and management of Kubernetes applications by:

**Packaging**: Bundling related Kubernetes resources together

**Templating**: Parameterizing manifests for reusability across environments

**Versioning**: Managing application versions and upgrades

**Release Management**: Tracking deployments and enabling rollbacks
{% endhint %}

1. Execute the following install script.

```bash
cd
cd ~/Pentaho-K3s-PostgreSQL/helm-chart
./install-helm.sh
```

<figure><img src="/files/nOWvRPyTvR55wXToMnQG" alt=""><figcaption><p>Install Helm</p></figcaption></figure>

or

1. Add the Helm GPG key.

```bash
# Add the Helm GPG key
curl https://baltocdn.com/helm/signing.asc | gpg --dearmor | sudo tee /usr/share/keyrings/helm.gpg > /dev/null# Install apt-transport-https package
```

2. Install dependencies

```bash
# Install apt-transport-https Package:
sudo apt-get install apt-transport-https --yes
```

3. Add Helm repository.

```bash
# Add the Helm repository
echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/helm.gpg] https://baltocdn.com/helm/stable/debian/ all main" | sudo tee /etc/apt/sources.list.d/helm-stable-debian.list

```

4. Update the package list and install.

```sh
# Update package list:
sudo apt-get update

# install Helm:
sudo apt-get install helm
```

5. Verify installation.

```sh
# Verify Helm Installation:
helm version --client
```

***

**Helm Commands**

```bash
# Checks your Helm chart for issues before deploying.
helm lint

# Renders your Helm chart templates locally and displays the output, useful for debugging.
helm template

# Simulates installing the chart without actually deploying it, allowing you to verify changes.
helm install my-chart . --dry-run

# Provides a list of all available Helm commands and options.
helm --help

# Updates an existing release with a new version of the chart, essential for maintaining and upgrading applications.
helm --help

# Lists all Helm releases in the current namespace, useful for tracking deployed applications.
helm list

# Retrieves the values used for a specific release, allowing you to review or modify configuration settings.
helm get values RELEASE_NAME

# Deletes a Helm release from the cluster, removing the associated Kubernetes resources.
helm delete RELEASE_NAME

# Rolls back a release to a previous version, useful for reverting to a stable state if an update causes issues.
helm rollback RELEASE_NAME REVISION

# Adds a Helm chart repository, enabling you to download charts from different sources.
helm repo add REPO_NAME REPO_URL

# Updates the local cache of chart repositories, ensuring you have the latest versions available.
helm repo update

# Packages a Helm chart into a .tgz file, which can be shared or uploaded to a repository.
helm package CHART_PATH

# Pushes a packaged chart to a Helm repository, facilitating distribution and deployment.
helm push CHART_PATH REPO_NAME
```

{% endtab %}

{% tab title="DBeaver" %}
{% hint style="info" %}

#### DBeaver

Your going to need a database management tool. DBeaver Community is a free, open-source database management tool for personal projects.
{% endhint %}

1. Simpliest option is to download & install from Snapstore.

```bash
cd
sudo apt update && sudo apt upgrade
sudo apt install snapd
sudo snap install dbeaver-ce
```

Or

Go to the official [DBeaver download page](https://dbeaver.io/download/)

{% embed url="<https://dbeaver.io/files/dbeaver-ce_latest_amd64.deb>" %}

Or

To install that DEB file.

```bash
wget https://dbeaver.io/debs/dbeaver.gpg.key -O /tmp/dbeaver.gpg.key
sudo gpg --dearmor -o /usr/share/keyrings/dbeaver.gpg /tmp/dbeaver.gpg.key
echo "deb [signed-by=/usr/share/keyrings/dbeaver.gpg] https://dbeaver.io/debs/dbeaver-ce /" | sudo tee /etc/apt/sources.list.d/dbeaver.list
sudo apt update
sudo apt install dbeaver-ce
```

4. Pin DBeaver to Dash - bottom toolbar.
   {% endtab %}

{% tab title="Make" %}
{% hint style="info" %}

#### Make

A makefile is simply a way of associating short names, called targets, with a series of commands to execute when the action is requested. For instance, a common makefile target is “clean,” which generally performs actions that clean up after the compiler - removing object files and the resulting executable.

We'll be using Make helper scripts to streamline the deployment process.
{% endhint %}

1. Update operating system.

```bash
sudo apt update && sudo apt upgrade
```

2. Check if make is already installed.

```bash
make -version
```

3. Install the make package.

```bash
sudo apt install make
```

4. Verify the installation.

```bash
ls /usr/bin/make
```

{% endtab %}
{% endtabs %}


# Common Questions

#### General

<details>

<summary>Where can I get a copy of the 'workshop files' ?</summary>

All the collateral can be found at: \~/Workshop--Installation.

You can also copy/fork the Git repository:

```
git clone https://github.com/jporeilly/Workshop--Installation
```

```git-rebase
gh repo clone jporeilly/Workshop--Installation
```

</details>

<details>

<summary>I cant write to sampledata database?</summary>

For security reasons the privileges from V10 have been removed. You will need to install Docker Desktop and create a MySQL sampledata database container.

</details>

<details>

<summary>Why do we need Docker &#x26; Docker Compose?</summary>

Its easier to manage some of the required applications - databases, Brokers, etc .. - as Docker containers.

</details>

#### Windows - Linux - MacOS (Self-paced Labs)

<details>

<summary>Where can i download 30-day Pentaho Enterprise Edition?</summary>

An 30-day activation code will be emailed to you:

<https://pentaho.com/download/#download-pentaho>

For BYOL:

<https://pentaho.com/pentaho-ee-onprem/>

</details>

#### Linux (Instructor-led Labs)

<details>

<summary>VM is unresponsive !</summary>

* Refresh the browser session to reconnect.
* Try another browser. The recommended browser: Google Chrome.
* If you're connecting via a Corporate VPN, then this may cause issues. Contact your IT dept to get the URL 'white' listed.

</details>

<details>

<summary>When does the Lab expire ?</summary>

The initial duration is 5 days. You will receive an email asking if you wish to extend your time limit.

</details>

<details>

<summary>Videos aren't loading?</summary>

* Hard refresh your browser. CTRL + F5

</details>

<details>

<summary>Is there sound ?</summary>

Yes .. There's no sound card attched to the Lab, so you'll need to copy and paste the Lab Guide URL in your host machine browser. 😊

</details>


# Pentaho Support

How to work with Pentaho Support and open effective tickets ..

{% hint style="info" %}
**Overview**

Pentaho Support at Pentaho helps you troubleshoot product issues and answer usage questions. Use this page to understand who can open tickets, what to include, and how to contact Support.
{% endhint %}

{% embed url="<https://www.youtube.com/watch?v=RRXw1d09RMk>" %}
Support Onboarding
{% endembed %}

<details>

<summary>Pentaho Virtualization Support Statement</summary>

Pentaho supports virtualization technology across the Pentaho Suite. This statement explains what is supported and what may be requested during troubleshooting.

#### What is supported

Pentaho products that run on:

* Supported operating systems
* Minimum hardware requirements

This applies whether you run Pentaho in a virtual environment or not.

#### What you are responsible for

You are responsible for failures caused by the hardware layer or operating system layer. This includes misuse of virtualization software.

#### What Support may ask you to do

Pentaho does not require you to reproduce every issue on physical hardware.

Pentaho may ask you to reproduce or diagnose an issue on a native (non-virtual) supported operating system.

This request is made only when there is reason to believe virtualization contributes to the issue.

#### How virtualization-related issues are handled

When a problem may relate to virtualization, investigation is handled as follows:

* Pentaho provides standard support for all products.
* If an issue occurs in a virtual environment, you may be asked to reproduce it on a physical (non-virtual) server. Pentaho then provides regular support.
* You can authorize Pentaho to investigate virtualization-related items at normal time and materials rates. If the issue is virtualization-related, you may request a software change, if a resolution is possible.
* If the issue is not virtualization-related, investigation and resolution are covered under regular maintenance.

Pentaho is expected to work in virtual environments.

#### Performance notes

Performance impacts may still occur. These impacts may or may not be caused by virtualization. They may fall outside this support statement.

</details>

<details>

<summary>Pentaho’s Global Data Protection &#x26; Privacy Policy</summary>

During troubleshooting, you may need to share data from your systems. This can include diagnostics, metadata, result sets, and similar artifacts.

#### Ways to send data

Pentaho Support provides several ways to transmit data, including:

* Email
* The Customer Portal
* Pentaho Content Platform (HCP)

#### Before you share

Share only what Support requests. Remove or redact sensitive data when possible.

See [Hitachi Vantara’s Global Data Protection & Privacy Policy](https://www.hitachivantara.com/en-us/legal) for details on how personal data and information are protected while in our possession.

</details>

{% tabs %}
{% tab title="Named Support Contacts" %}
{% hint style="info" %}
Pentaho Support works directly with your organization’s **Named Support Contacts**. These are specific people with a **unique (non-generic) email address** and a **phone number**.
{% endhint %}

#### Requirements and responsibilities

* **Language and ticket ownership**\
  Contacts must communicate in English. They own ticket updates and follow-ups.
* **Portal access and ticket submission**\
  Named Support Contacts can access the Customer Support Portal.\
  Only **Primary** and **Backup** contacts can open new tickets.
* **Administrative access**\
  Troubleshooting often requires admin access or elevated permissions.\
  Primary contacts should be able to approve or perform sensitive actions.
* **Keep contacts current**\
  Tell Support when a contact leaves or changes roles.\
  This ensures access is revoked and tickets are reassigned.
* **Single point of contact**\
  Named Support Contacts should not forward requests from others.\
  They act as the Support liaison for their organization.
* **Service announcements**\
  Primary contacts automatically receive critical announcements.

#### What each access level can do

| Task                    | Primary and Backup | Other Named Support Contacts | Anonymous |
| ----------------------- | ------------------ | ---------------------------- | --------- |
| Submit a ticket         | Yes                | No                           | No        |
| Knowledge Base          | Yes                | Yes                          | Limited   |
| Best practices          | Yes                | Yes                          | Yes       |
| Downloads               | Yes                | Yes                          | No        |
| Product documentation\* | Yes                | Yes                          | Yes       |
| Academy                 | Yes                | Yes                          | Yes       |

{% hint style="warning" %}
If you are licensed for this product but you do not have access, email **<support.pentaho@pentaho.com>**.
{% endhint %}
{% endtab %}

{% tab title="Open a ticket" %}
{% hint style="info" %}
Ticket handling is guided by three factors:

* **Subscription service level**\
  Defined in your subscription agreement.\
  If you are unsure, contact your Customer Success Manager or Sales Representative.
* **Severity level**\
  Based on business impact and user impact.
* **Environment**\
  Where the issue occurred (Production, Development, Test).

When you open a ticket, you’ll specify **severity** and **environment**. Your subscription service level is identified automatically.
{% endhint %}

#### Severity definitions

Use severity to describe **business impact** and **user impact**. Support may adjust severity based on the details provided.

* **Severity 1 (Critical)**\
  Production is down or a critical function is unavailable.\
  No workaround exists. Impact is immediate and widespread.
* **Severity 2 (High)**\
  Major functionality is impaired or performance is severely degraded.\
  A workaround exists, but it is limited or risky. Impact is significant.
* **Severity 3 (Medium)**\
  Minor functionality is affected, or you have a how-to question.\
  A workaround exists and impact is limited.

{% hint style="info" %}
For Premium and Enterprise customers, Severity 1 issues in Production may be eligible for after-hours coverage.
{% endhint %}

{% stepper %}
{% step %}

#### Do a quick self-check first

* Confirm the issue is in Pentaho, not third-party software.
* Confirm you’re on a supported version. See [Pentaho EOL Policy](https://support.pentaho.com/hc/en-us/articles/205789159-Pentaho-Product-Lifecycle-Overview).
* Try to reproduce the issue. Note if it’s consistent or intermittent.
* Search for known solutions in the Knowledge Base.

Useful resources:

* [Pentaho Product Documentation](https://docs.pentaho.com/)
* [Pentaho Academy](https://academy.pentaho.com/)
  {% endstep %}

{% step %}

#### Gather the right details

If you’re **asking a question**, include:

* What you’re trying to achieve.
* Your Pentaho product(s) and version(s).
* Environment details (OS, JVM, DB, cloud/on-prem, and so on).
* What you already tried, including exact URLs to docs or KB articles.

If you’re **reporting a product issue**, also include:

* Steps to reproduce (numbered).
* Expected result vs actual result.
* Symptoms and exact error messages.
* Logs and diagnostics. Use the [Pentaho Support Utility](/pentaho-11-installation-en/pentaho-support/pentaho-support/pentaho-support-utility) when possible.
* What changed before it started (deployments, config, certificates, upgrades).
* When it started and how often it happens.
* Business impact and urgency.
  {% endstep %}

{% step %}

#### Submit the ticket (Portal or email)

Only **Primary** and **Backup** Named Support Contacts can open new tickets.

#### Submit via portal

Use the [Pentaho Customer Portal](https://support.pentaho.com/hc/en-us) to submit a request and open a ticket.

<figure><img src="/files/nHbFI3Tc1vwP9VuWIgmd" alt=""><figcaption><p>Submit via Portal</p></figcaption></figure>

#### Submit via email

Email **<support.pentaho@hitachivantara.com>** with the details above.\
If you are a Named Support Contact, you’ll receive a confirmation email with a ticket number.

To add more information later, reply to the ticket emails.\
Include the ticket number in the email subject.

<figure><img src="/files/cp5zkGZnkcSMtbrKsaCN" alt=""><figcaption><p>Subject Line</p></figcaption></figure>
{% endstep %}

{% step %}

#### Attach visuals when it helps

Screenshots and short videos can speed up triage and reproduction. Blur or redact sensitive data before sharing.
{% endstep %}

{% step %}

#### Assigned to the support team

Support typically triages and resolves tickets in levels. If needed, Support escalates to the next level.

| Support team level                                                | Description                                                                                                                                                                         |
| ----------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Level 0: Self-help and user-retrieved information**             | Self-service resources such as FAQs, Knowledge Base articles, forums, product documentation, and troubleshooting guides. Users resolve issues independently.                        |
| **Level 1: Basic help desk resolution and service desk delivery** | Tier 1 frontline Support handles incoming inquiries, troubleshoots common problems, and provides known fixes and standard workarounds.                                              |
| **Level 2: Advanced technical support**                           | Tier 2 handles issues that Level 1 cannot resolve. This often requires deeper product knowledge and more advanced troubleshooting. Tickets may be escalated to Level 3.             |
| **Level 3: Expert product and service support**                   | Tier 3 specialists handle complex or product-specific issues. They may work with engineering on defects, enhancements, or unique requirements. Tickets may be escalated to Level 4. |
| **Level 4: Product and engineering development**                  | Engineering investigates and implements code-level fixes for the most complex issues. This level is typically involved when product changes are required.                           |
| {% endstep %}                                                     |                                                                                                                                                                                     |
| {% endstepper %}                                                  |                                                                                                                                                                                     |

#### What happens after you submit a ticket

**During business hours**

* Tickets are assigned to the next available engineer for the product area.
* The Support engineer may call you to confirm details and start quickly.
* If you cannot be reached by phone, Support contacts you in the ticket.

**Outside business hours and global holidays**

* Tickets submitted after local business hours are picked up the next business day.
* Tickets submitted on global holidays are picked up the next business day.
* For **Premium or Enterprise Severity 1 issues**, the on-call engineer responds after hours.
  {% endtab %}

{% tab title="Support hours" %}

#### Support hours

* Standard coverage: **9:00 AM to 5:00 PM, Monday to Friday** (local Support office hours).
* **24/7/365** support is available for **Premium and Enterprise** customers for **Severity 1 issues in Production**. It can also be arranged in advance for events like migrations or go-live dates.

#### Response times

Initial response targets depend on:

* Your **subscription service level**
* The ticket **severity**
* The affected **environment** (Production vs non-Production)

| Severity Level          | Enterprise       | Premium          | Starter & Standard |
| ----------------------- | ---------------- | ---------------- | ------------------ |
| **24/7/365 Production** | Yes              | Yes              | No                 |
| Severity 1              | 1 business hour  | 1 business hour  | 4 business hours   |
| Severity 2              | 2 business hours | 2 business hours | 1 business day     |
| Severity 3              | 4 business hours | 4 business hours | 2 business days    |

{% hint style="info" %}
Response-time targets vary by contract.\
If you need your specific SLA, contact your Customer Success Manager or ask Support in your ticket.

Enterprise Support includes up to **4 hours per week** of Customer Success time (Solution Architects, Customer Success Managers (CSM), or Support Account Managers (SAM)).

This time can be used to:

* Run best-practice sessions with your Named Support Contacts.
* Discuss integration techniques, solution design, and architecture.
* Discuss implementation strategies and upgrade approaches.
* Review best practices and performance tuning.
* Coordinate sessions with Pentaho subject matter experts.
* Troubleshoot issues in your systems or create solution replicas, when technically possible.

Weekly time does not accrue or roll over. You can increase the allocation by written agreement.
{% endhint %}
{% endtab %}

{% tab title="Escalate a ticket" %}
{% hint style="info" %}

#### Escalate a ticket

If the ticket priority has changed or expectations are not met, a **Primary contact** can escalate by emailing **<escalation.pentaho@hitachivantara.com>**. Include the ticket number in the subject line.

The ticket owner continues working with you on impact and next steps. Support also opens an escalation ticket with a Support management team member.
{% endhint %}
{% endtab %}

{% tab title="Defects & Enhancement Requests" %}
{% hint style="info" %}

#### Defects and enhancement requests

If Support reproduces the issue and suspects a defect, the support engineer links your ticket to a Jira issue in Product Engineering. Support changes the ticket status to **Defect** and shares major updates. Your ticket stays open while the defect is investigated.

If the request is an enhancement, Support may create a Jira enhancement. Enhancements do not have a committed implementation timeline. After requirements are captured, Support closes the ticket. Open a new ticket later if you need a status update.
{% endhint %}

<figure><img src="/files/d0m1HbuaeCmCjNHd31Uo" alt=""><figcaption><p>Ticket Status</p></figcaption></figure>

#### Automated closure timeline

| Timeframe                         | What happens                                                                                                                 |
| --------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| 24 hours after status **Pending** | You get an email that Support needs your response.                                                                           |
| 7 days after status **Pending**   | You get an email that Support has waited 7 days. Support may consider the issue resolved. Reply to keep the ticket **Open**. |
| 14 days after status **Pending**  | The system sets the ticket to **Solved** and says it will close in 7 days. A feedback survey is emailed within 24 hours.     |
| 21 days after status **Pending**  | The system sets the ticket to **Closed**. Closed tickets cannot be reopened. Submit a new ticket.                            |

{% hint style="info" %}
The automated closure process is suspended for escalated tickets and for tickets linked to a defect. You also receive a final email at ticket closure inviting you to complete a customer satisfaction survey.
{% endhint %}

{% hint style="warning" %}

#### Exclusions from Support

Pentaho does not provide support for issues caused by:

* Hardware, equipment, or programs not covered by your contract.
* Product versions not obtained through Pentaho Support.
* Use of non-general availability releases in Production (not marked as general availability (GA)).
* Causes outside Pentaho’s control (for example, floods, fires, or power loss).
* Failures outside Pentaho software (for example, databases, web servers, or hardware).
* Not following operating instructions in product documentation.
* Modifications, enhancements, or customizations performed by anyone outside Pentaho.
* Installation, configuration, management, and operation of your applications.
* APIs, interfaces, web services, or data formats not included with the product.
* Third-party products, except those provided by Pentaho, and those used only to support intended interfaces or functionality.
  {% endhint %}
  {% endtab %}
  {% endtabs %}


# Pentaho Support Utility

{% hint style="info" %}
The Pentaho Support Utility collects environment and configuration details into a single zip file that you can attach to your support ticket. This gives the support team immediate access to the background information they need to diagnose and resolve your issue, reducing back-and-forth requests for additional files and minimizing the time you spend gathering data.

The utility is supported on Pentaho versions 7.x, 8.x, 9.x, and 10.x. For a complete list of supported software and hardware, refer to the Components Reference in Pentaho Documentation.

Three tools are available for running the utility:

**Pentaho Server Plugin**: provides a UI within the Pentaho Server to gather information from the currently running server instance.

**Pentaho Data Integration Plugin**: operates within Spoon to collect details from your PDI environment.

**Command Line Utility**: offers a standalone option for gathering information from either the Pentaho Server or PDI without requiring a running application.

The server and PDI plugins are the preferred methods for running the utility. Use the command line tool only when Pentaho fails to start or when organizational policies prohibit plugin installation.

If any individual collector fails during execution, the remaining collectors will continue to run and the resulting report will still provide value to the support team. You can identify failed collectors by reviewing the utility logs or checking for .failed files within the output zip.

Some collected files and configuration details may contain passwords. While the utility attempts to remove all passwords during collection, Pentaho cannot guarantee complete removal in every case - particularly for files such as kettle.properties.
{% endhint %}

1. Download the Pentaho Support Utility.

{% embed url="<https://hcpanywhere.hitachivantara.com/userportal/?v=4.6.1#/shared/public/8eUdzjLlHa7j5R3R/245ce549-c03a-4ee7-adae-d5ea655af8ad>" %}

<figure><img src="/files/nAf9rzfIwbfXSv0e1COl" alt=""><figcaption><p>Support Utilities</p></figcaption></figure>

2. Select Pentaho Product:

{% tabs %}
{% tab title="Pentaho Server" %}
**To install the Pentaho Server Plugin:**

1. Download: `pentaho-support-utility-server-plugin.zip` file.
2. Unzip this file to: `pentaho-server/pentaho-solutions/system`.

```bash
cd
cd ~/Downloads/'Support Utility'/
unzip pentaho-support-utility-server-plugin-1.0.1.zip -d /opt/pentaho/server/pentaho-server/pentaho-solutions/system
```

{% hint style="info" %}
If needed, configure the [password removal regular expression](https://support.pentaho.com/hc/en-us/articles/360021624372-How-to-Install-and-Use-Pentaho-Support-Utility#PRCF).
{% endhint %}

3. Restart the Pentaho Server.

```sh
cd
cd /opt/pentaho/server/pentaho-server
./stop-pentaho.sh
```

```sh
cd
cd /opt/pentaho/server/pentaho-server
./start-pentaho.sh
```

{% endtab %}
{% endtabs %}


# Pentaho 10 Installation

Archive Installation of Pentaho Pro Platform ..

### Introduction

{% hint style="info" %}
Pentaho Enterprise is a data integration and analytics platform that offers various tools and features for data ingestion, transformation, visualization, and reporting. Pentaho Pro can run on different operating systems, including Linux.

The following three production installation methods are available:

* [**Archive**](/pentaho-10-installation/installation/archive-installation)

Choose this option if you want to run the Pentaho Server on the version of Tomcat which we supply.

* **Manual**

Choose this option if you want to deploy the Pentaho Server on your existing Tomcat or JBoss web app server.

* [**Client Tools**](/pentaho-10-installation/installation/archive-installation/install-client-tools)

Choose this option if you want to install Business Analytics (BA) or Data Integration (DI) components only.

This workshop walks you through an 'Archive' installation of the Pentaho Client / Server on a Tomcat web application server supplied by Pentaho; with a default PostgrSQL 15 Pentaho Repository database.

By the end of the installation, you will have a comprehensive understanding of:
{% endhint %}

<details>

<summary><a href="/pages/2DAbLXeZyinGtWELgg1R">Pentaho Enterprise components</a></summary>

If your just starting your Pentaho journey, then dive into the components that comprise the Pentaho Enterprise Edition.

The core architecture consists of several key layers: the Pentaho Server (also known as the BI Platform), which serves as the central hub for managing users, security, scheduling, and content repository; the Pentaho Data Integration (PDI) engine for ETL processes; and various client tools and interfaces that connect to these services.

</details>

<details>

<summary><a href="/pages/cTu381Tu5j7GsbWNNn90">Linux Installation</a></summary>

The basic installation steps include extracting the archive to a designated directory (commonly `/opt/pentaho`), setting up a supported database like PostgreSQL or MySQL for the Pentaho repository, and running the provided database initialization scripts. The installer requires configuring environment variables, particularly `PENTAHO_HOME` and Java settings, as the platform runs on Java and requires a compatible JDK version.

</details>

<details>

<summary><a href="/pages/zEi0WmIdEphF6qihuEvC">Post Installation Tasks</a></summary>

After the initial setup, administrators need to configure the server properties, including database connections, security settings, and port configurations. The default installation runs Pentaho Server on port 8080 and includes the Pentaho User Console (PUC) for web-based administration and report management.

</details>

<details>

<summary><a href="/pages/6xTrTpUZvWJMRkBkl3o8">Upgrades &#x26; Patches</a></summary>

Pentaho upgrades and patches involve a systematic process to enhance functionality, security, and performance of the business intelligence platform. The upgrade process typically begins with thorough planning, including backing up existing data, configurations, and custom components.

Organizations must evaluate their current version, review release notes for new features and breaking changes, and test the upgrade in a development environment before applying it to production systems.

</details>

<details>

<summary><a href="/pages/Y2pnhKgldYUSoHG9Pane">Windows Installation</a></summary>

This evaluation wizard setup allows organizations to assess Pentaho's suitability for their business intelligence needs, test integration with existing data sources, and evaluate the user experience before committing to a full commercial license. The installation serves as a sandbox environment for exploring the platform's analytics and reporting capabilities within a Windows infrastructure.

</details>

***


# Pentaho Support

Take a look at the Welcome Letter ..

{% file src="/files/1sWSBeqdmHWHV7tNPdlw" %}

Support Onboarding ..

{% embed url="<https://www.youtube.com/watch?v=RRXw1d09RMk>" %}
Support Onboarding
{% endembed %}

{% tabs %}
{% tab title="Tips for Submitting a Support Ticket" %}
{% hint style="info" %}
In order to assist us in addressing your inquiry or concern promptly and effectively, please be prepared to provide the following information:
{% endhint %}

{% stepper %}
{% step %}
**Explore Pentaho Documentation & Academy**

Great resources with an AI assistent to help locate the information:

[Pentaho Product Documentation](https://docs.pentaho.com/)

[Pentaho Academy](https://academy.pentaho.com/)
{% endstep %}

{% step %}
**What information is required?**

If **asking a question**, please provide:

* A description of what you are trying to achieve.
* Configuration details regarding the Pentaho Suite and the affected environment/version.
* If applicable, outline the steps you have already taken to either:
  * Consult support articles or the [Pentaho Product Documentation](https://docs.pentaho.com/) - please provide the exact URL.
  * Attempt to replicate the scenario.

If **reporting an issue** with the product, in addition to the above, please provide:

* Symptoms experienced and, if applicable, the exact error message(s).
* Log files or other supporting documentation - For Pentaho, we recommend utilizing our [Support Utility](https://support.pentaho.com/hc/en-us/articles/360021624372-How-to-Install-and-Use-Pentaho-Support-Utility).
* What changes have occurred that may have led to this condition?
* Does the condition occur intermittently or every time a specific action is executed?
* When did the condition first manifest?
* If applicable, what steps have you taken to isolate and/or resolve the issue? What were the outcomes?
* Impact on your business operations.
  {% endstep %}

{% step %}
**A screenshot or video can be worth a thousand words…**

While detailed descriptions are incredibly useful for troubleshooting or addressing your question, a screenshot or video can be even more effective in helping us understand or replicate your issue.
{% endstep %}
{% endstepper %}
{% endtab %}

{% tab title="Pentaho Support Utility" %}
{% hint style="info" %}
The Pentaho Support Utility collects environment and configuration details into a single zip file that you can attach to your support ticket. This gives the support team immediate access to the background information they need to diagnose and resolve your issue, reducing back-and-forth requests for additional files and minimizing the time you spend gathering data.

The utility is supported on Pentaho versions 7.x, 8.x, 9.x, and 10.x. For a complete list of supported software and hardware, refer to the Components Reference in Pentaho Documentation.

Three tools are available for running the utility:

**Pentaho Server Plugin**: provides a UI within the Pentaho Server to gather information from the currently running server instance.

**Pentaho Data Integration Plugin**: operates within Spoon to collect details from your PDI environment.

**Command Line Utility**: offers a standalone option for gathering information from either the Pentaho Server or PDI without requiring a running application.

The server and PDI plugins are the preferred methods for running the utility. Use the command line tool only when Pentaho fails to start or when organizational policies prohibit plugin installation.

If any individual collector fails during execution, the remaining collectors will continue to run and the resulting report will still provide value to the support team. You can identify failed collectors by reviewing the utility logs or checking for .failed files within the output zip.

Some collected files and configuration details may contain passwords. While the utility attempts to remove all passwords during collection, Pentaho cannot guarantee complete removal in every case - particularly for files such as kettle.properties.
{% endhint %}

1. Download the Pentaho Support Utility.

{% embed url="<https://hcpanywhere.hitachivantara.com/userportal/?v=4.6.1#/shared/public/8eUdzjLlHa7j5R3R/245ce549-c03a-4ee7-adae-d5ea655af8ad>" %}

<figure><img src="/files/nAf9rzfIwbfXSv0e1COl" alt=""><figcaption><p>Support Utilities</p></figcaption></figure>

2. Select Pentaho Product:

{% tabs %}
{% tab title="Pentaho Server" %}
**To install the Pentaho Server Plugin:**

1. Download: `pentaho-support-utility-server-plugin.zip` file.
2. Unzip this file to: `pentaho-server/pentaho-solutions/system`.

```bash
cd
cd ~/Downloads/'Support Utility'/
unzip pentaho-support-utility-server-plugin-1.0.1.zip -d /opt/pentaho/server/pentaho-server/pentaho-solutions/system
```

{% hint style="info" %}
If needed, configure the [password removal regular expression](https://support.pentaho.com/hc/en-us/articles/360021624372-How-to-Install-and-Use-Pentaho-Support-Utility#PRCF).
{% endhint %}

3. Restart the Pentaho Server.

```sh
cd
cd /opt/pentaho/server/pentaho-server
./stop-pentaho.sh
```

```sh
cd
cd /opt/pentaho/server/pentaho-server
./start-pentaho.sh
```

4.

x

x
{% endtab %}

{% tab title="Pentaho Data Integration" %}
x

x

x
{% endtab %}

{% tab title="Command Line" %}
x

x

x

x
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}

x


# Getting Started

Archive installation of Pentaho Enterprise on Linux - Ubuntu 22.04 LTS ..


# Components

Overview of Pentaho Pro components ..

{% hint style="info" %}

#### Pentaho Client / Server Architecture

Pentaho's client/server architecture forms the basis of its data integration and business analytics suite, providing a flexible and scalable platform for enterprise data management and analysis. The architecture is designed to support various data integration, reporting, and analytics needs across an organization.
{% endhint %}

<figure><img src="/files/AanhZpaBxQDvyJLN3kGj" alt=""><figcaption><p>Client / Server</p></figcaption></figure>

<table><thead><tr><th width="186">Port Number</th><th>Description</th></tr></thead><tbody><tr><td>5432</td><td>PostgreSQL Server</td></tr><tr><td>8080</td><td>Pentaho Server Tomcat Web Server Startup Port</td></tr><tr><td>8012</td><td>Pentaho Server Shutdown Port</td></tr><tr><td>9001</td><td>HSQL Server Port</td></tr><tr><td>9092</td><td>Embedded H2 Database</td></tr></tbody></table>

Key components include:

{% tabs %}
{% tab title="Pentaho Client Tools" %}
{% hint style="info" %}

#### Pentaho Client Tools

The Pentaho Client, a key component of the Pentaho suite, encompasses several user-facing tools designed for data management and analytics. These include the Data Integration tool (PDI), which is central to extracting, transforming, and loading (ETL) operations; Spoon, a graphical user interface for designing ETL processes; Designer for convenient pipeline design; Scheduler linked to Quartz for job scheduling; Repository Browser for managing ETL assets; and Database Explorer for database operations.

Additionally, it offers tools like Metadata Editor and Schema Workbench for advanced data manipulation. Together, these tools empower users to efficiently process and analyze data within the Pentaho ecosystem.
{% endhint %}

{% tabs %}
{% tab title="Data Integration" %}
{% hint style="info" %}

#### Data Integration

Pentaho Data Integration (PDI), also known as Kettle, is an open-source data integration tool that allows the extraction, transformation, and loading (ETL) of data into databases, data warehouses, and business applications. It is designed to handle a wide variety of data sources including traditional relational databases, unstructured data formats, and cloud-based storage. PDI is composed of several key components that work together to provide a comprehensive ETL solution.
{% endhint %}

<figure><img src="/files/kSv68yqrodeKvgOZaTT8" alt=""><figcaption><p>Pentaho Client / Server Architecture</p></figcaption></figure>

{% hint style="info" %}

#### Spoon

Spoon is the graphical user interface (GUI) for designing and testing PDI jobs and transformations. It allows users to visually create, edit, and manage ETL processes without writing code.

#### Designer

Drag & Drop 'objects' to design your pipelines and workflows.

#### Scheduler

Connects to Quartz scheduler on server. Jobs and transformations must be uploaded to Repository.

#### Repository Browser

The repository is a central storage area for PDI resources such as jobs, transformations, and database connections. It facilitates collaboration among team members by allowing them to share and manage ETL assets efficiently.

These components collectively make PDI a powerful tool for data integration, enabling businesses to cleanse, integrate, and analyze data from diverse sources more effectively.

Connects to Apache Jackrabbit content Repository, pointing to a supported database:

* PostgreSQL
* MSSQL Server
* Oracle
* MySQL
* MariaDB

#### DB Explorer

Database Explorer that enables you to conduct minimal database operations.
{% endhint %}
{% endtab %}

{% tab title="Metadata Editor" %}
{% hint style="info" %}

#### Metadata Editor

The Pentaho Metadata Editor is a tool within the Pentaho suite that facilitates the creation and management of business models. These models form the foundation for reporting and analysis, making it easier for end-users to interact with data without needing a deep understanding of the underlying database structures.

Key features include:

**User-friendly Interface:** Offers a graphical environment where users can define business models, relationships, and metadata concepts, simplifying complex data structures into more understandable terms.

**Data Source Connection:** Allows connection to various data sources, enabling the extraction of metadata from relational databases, OLAP sources, and more.

**Security Settings:** Supports the definition of security constraints at the model level, ensuring that sensitive data remains protected and access is controlled.

**Localization and Internationalization:** Models can be localized, allowing the presentation of metadata in different languages to support global deployments.

The Metadata Editor plays a crucial role in the Pentaho Business Analytics suite, streamlining the creation of complex reports and analyses by offering a simplified view of data for business users.
{% endhint %}

<figure><img src="/files/Vx6d3sfOqf0QsLOrrj2m" alt=""><figcaption><p>Metadata Editor</p></figcaption></figure>
{% endtab %}

{% tab title="Schema Workbench" %}
{% hint style="info" %}

#### Schema Workbench

The Pentaho Schema Workbench is an essential tool within the Pentaho suite designed for developers and data architects to create and edit OLAP (Online Analytical Processing) schemas. It provides a graphical interface for defining the multidimensional models needed for complex analytical queries, enabling the efficient organization and visualization of large data sets.

With its user-friendly interface, users can easily design OLAP cubes that form the foundation of advanced analytics and business intelligence applications, making data more actionable and insights more accessible.
{% endhint %}

<figure><img src="/files/eOwrGYRxkKJ5rfWk2bTA" alt=""><figcaption><p>Schema Workbench</p></figcaption></figure>
{% endtab %}

{% tab title="Aggregation Designer" %}
{% hint style="info" %}

#### Aggregation Designer

The Pentaho Aggregation Designer is a pivotal tool aimed at improving query performance by simplifying the creation and management of aggregate tables in a star schema database. This graphical tool assists users in defining, generating, and deploying SQL-based aggregation tables that summarily condense detailed data into summarized formats, making data retrieval processes significantly more efficient for analytical queries.

This capability is critical for enhancing the performance of OLAP cubes, facilitating faster data analysis, and providing a more streamlined user experience in the Pentaho Business Analytics suite.
{% endhint %}

<figure><img src="/files/9NN78vsfHz4NK9OzCMeW" alt=""><figcaption><p>Aggregation Designer</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Pentaho Server" %}
{% hint style="info" %}

#### Pentaho Server

Pentaho Server acts as the central platform for hosting and managing all Pentaho applications and services. It provides a secure, scalable environment for deploying and executing Pentaho's analytics and data integration solutions. Key components include:

* **BI Server:** Facilitates interactive reporting, analytics, dashboarding, and data exploration.
* **Data Integration Server:** Supports the orchestration and scheduling of ETL (Extract, Transform, Load) processes.
* **User Console:** Offers a web-based interface for accessing, creating, and managing content within the Pentaho suite.
* **Security:** Integrates with enterprise security systems to provide authentication, authorization, and secure access.
* **Repository:** Centralizes the storage of all Pentaho assets, including reports, dashboards, and ETL scripts, ensuring collaboration and version control.

The server enables organizations to leverage the full potential of the Pentaho suite by providing a comprehensive platform for business intelligence and data management activities.
{% endhint %}

#### Pentaho Server Reporting Suite

{% tabs %}
{% tab title="Analyzer" %}
{% hint style="info" %}

#### Analyzer

Pentaho Analyzer is an interactive analytics and data visualization tool that is part of the Pentaho Business Analytics suite. It enables users to explore and analyze data through an intuitive web-based interface, providing rich graphical representations of data including charts, tables, and heat maps. Users can create and customize reports and dashboards without the need for in-depth technical knowledge, making it accessible to a wide range of users. Key features include:

* **Ad-hoc analysis:** Empowers users to quickly create and modify reports based on their specific questions and needs.
* **Drag-and-drop interface:** Simplifies the process of designing reports by allowing users to easily select and arrange data elements.
* **Rich visualizations:** Supports a wide array of visualization options to help users uncover insights from their data.
* **Collaboration and sharing:** Enables sharing of reports and dashboards with other users to facilitate decision-making across teams and departments.

Pentaho Analyzer is designed to work seamlessly with the Pentaho suite, integrating directly with Pentaho's data integration, ETL, and data warehousing capabilities. This allows users to leverage the full power of the suite for comprehensive data analysis and business intelligence solutions.
{% endhint %}

<figure><img src="/files/cVCAZwpJBuC21Vj660JP" alt=""><figcaption><p>Analyzer Report</p></figcaption></figure>
{% endtab %}

{% tab title="Interactive Reports" %}
{% hint style="info" %}

#### Interactive Reports

Pentaho Interactive Reports offer a highly user-friendly interface for creating, editing, and viewing ad-hoc reports. This feature is designed for business users who need to generate reports quickly without in-depth technical knowledge of the underlying data structure.

* **User-Friendly Interface:** Provides a drag-and-drop interface, making it easy for users to select, organize, and present data without any SQL knowledge.
* **Real-Time Data Exploration:** Enables users to interact with their data in real-time, allowing for instant filtering, sorting, and aggregation to identify trends and insights.
* **Customizable Layouts:** Users can customize the layout of their reports by adjusting columns, rows, and summaries to meet their specific reporting needs.
* **Export and Share:** Reports can be exported to various formats (e.g., PDF, Excel, CSV) and shared with stakeholders to support data-driven decision-making.

Interactive Reports are part of the larger Pentaho Business Analytics suite, offering seamless integration with Pentaho's ETL and data analysis tools, ensuring businesses have a comprehensive solution for their data integration and reporting needs.
{% endhint %}

<figure><img src="/files/HEMnTvFSEC5Qvohg56fL" alt=""><figcaption><p>Interactive Report</p></figcaption></figure>
{% endtab %}

{% tab title="Dashboard Designer" %}
{% hint style="info" %}

#### Dashboard Designer

Pentaho Dashboard Designer is a feature-rich tool within the Pentaho Business Analytics suite, designed for creating interactive and visually appealing dashboards. These dashboards aggregate and display data from various sources, providing users with insights at a glance. Here's a quick overview:

* **Intuitive Design Interface**: Offers a drag-and-drop interface, making it accessible for non-technical users to create and customize dashboards.
* **Data Integration**: Seamlessly integrates with Pentaho Data Integration (PDI), allowing it to pull data from a wide range of sources for real-time analytics.
* **Interactive Widgets**: Supports various types of widgets including charts, tables, and filters, enabling interactive data exploration.
* **Customization and Branding**: Allows for the customization of layout and design, enabling alignment with company branding.
* **Collaboration Features**: Facilitates sharing and collaboration by allowing users to publish dashboards within the organization or to a broader audience.
* **Security**: Integrates with existing security frameworks, ensuring data protection and controlled access based on roles and permissions.

Pentaho Dashboard Designer plays a crucial role in transforming data into actionable insights, driving informed decision-making across organizations.
{% endhint %}

<figure><img src="/files/nJqBoUswZk516r8tYuS5" alt=""><figcaption><p>Dashboard</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Carte Server" %}
{% hint style="info" %}

#### Carte Server

Pentaho Carte is a lightweight web server for remote execution and monitoring of ETL processes created in Pentaho Data Integration (PDI/Kettle).

Carte is built on Java and uses the embedded Jetty web server. It relies on XML-based configuration and exposes functionality through a REST API, with a simple browser-based interface for monitoring.

The server enables remote execution of transformations and jobs, supports clustering for load balancing, provides real-time monitoring, and allows scheduling of ETL processes.

Carte can be deployed as a standalone server, in a master-slave cluster setup, or in a load-balanced environment for high availability. It's typically launched via command line with a configuration file containing server settings.

This component is crucial for Pentaho's distributed processing architecture, allowing organizations to scale data integration processes across multiple machines.
{% endhint %}

<figure><img src="/files/lkQgwcFsbqVev1cEb6M0" alt=""><figcaption><p>Carte Cluster</p></figcaption></figure>
{% endtab %}

{% tab title="APIs" %}
{% hint style="info" %}

#### Kitchen

Kitchen is a command-line tool that enables the execution of PDI jobs. It supports batch processing and can be integrated into automated workflows, allowing for efficient data processing.
{% endhint %}

```
kitchen.sh -file=/PRD/updateWarehouse.kjb -level=Minimal
kitchen.bat /file:D:\Jobs\updateWarehouse.kjb /level:Basic
```

{% hint style="info" %}

#### Pan

Similar to Kitchen, Pan is a command-line tool but is specifically designed for executing PDI transformations. It provides flexibility in running ETL transformations from shell scripts or scheduling systems.
{% endhint %}

```
pan.sh -file="/PRD/Customer Dimension.ktr" -level=Minimal
pan.bat /file:"D:\Transformations\Customer Dimension.ktr" /level:Basic
```

{% embed url="<https://docs.pentaho.com/pentaho-rest-api>" %}
{% endtab %}
{% endtabs %}


# Linux Installation

Recommended installation option ..

{% hint style="info" %}

#### Introduction

The workshop walks you through installation of the Pentaho Server on a Tomcat web application server supplied by Pentaho and configuration of the Pentaho Repository database of your choice.

The archive installation method provides you with a preconfigured Tomcat web application server. This method is most useful when you do not already have a web app server. As such, the archive installation method is an easier and quicker installation process versus a manual installation. This method is an archive by virtue of its being an image or "snapshot" of a Pentaho configuration built on Tomcat.

An archive installation is also useful if you already have existing Pentaho data or repositories that you created with a previous version of Pentaho or that you created with an evaluation installation of Pentaho. The archive installation method is an easy way to bring your existing Pentaho data into your production version.
{% endhint %}

* [x] **Before You Begin**:
  * Read the overview of installation options in the [Pentaho Installation article](https://docs.pentaho.com/pdia-10.2-install/pentaho-installation-overview-cp) to ensure that this method is the best installation option for you.
  * Check the [Components Reference](https://docs.pentaho.com/pdia-10.2-install/components-reference) to verify that your server computer, Pentaho Repository database, and web browser meet Pentaho’s requirements for this version of the software.
  * Uninstall any evaluation version of Pentaho.
* [x] **Requirements**:
  * You’ll need a server device that meets the hardware and software requirements specified in the Components Reference.
  * Ensure you have either the Oracle Java Runtime Environment (JRE) or Oracle Java Development Kit (JDK) installed.
  * Choose a Pentaho Repository database (PostgreSQL, MySQL, MS SQL Server, or Oracle). You’ll need to supply, install, and configure your chosen database yourself.
* [x] **Installation Steps**:
  * **Linux Environment**:
    * Create the 'Pentaho / Install' user - ensure has the required permissions.
    * Set up the Linux directory structure.
    * Install Java.
    * Install the Pentaho Repository host database.
    * Download and unpack the installation files.
    * Set environment variables. Detailed instructions can be found [here](https://help.hitachivantara.com/Documentation/Pentaho/8.2/Setup/Installation/Archive/Linux_Environment).
  * **Windows Environment**:
    * Create the Windows directory structure.
    * Install Java.
    * Install the database that will host the Pentaho Repository.
    * Download and unpack the installation files. More details are available [here](https://help.hitachivantara.com/Documentation/Pentaho/Data_Integration_and_Analytics/9.0/Setup/Prepare_your_Windows_environment_for_an_archive_install).
* [x] **Start the Pentaho Server**:
  * [After completing the above steps, start the Pentaho Server and install your licenses](https://docs.pentaho.com/pdia-10.2-install/pentaho-installation-overview-cp/archive-installation/archive-installation-process/starting-the-pentaho-server-after-an-archive-installation).


# Prepare Environment

Preflight tasks ..

{% hint style="info" %}

#### Prepare Environment

You will need perform the following steps to prepare your Linux environment for an Archive installation of the Pentaho Server.

This process includes:

* create a 'Pentaho installation user' with sudo privileges.
* install supported version of OpenJDK 11
* set PENTAHO\_JAVA\_HOME
* install certified version of Postgres 15
* create 'superadmin' user
* install pgAdmin4
  {% endhint %}

{% hint style="warning" %}
The supported Linux environment for Pentaho version 10.2.x: Ubuntu 22.04

Check [Components Reference](https://docs.pentaho.com/pdia-10.2-install/components-reference)
{% endhint %}

{% tabs %}
{% tab title="1. Creating a Pentaho User" %}
{% hint style="info" %}

#### Pentaho Installation Account

Add an account that is assigned administrative privileges by performing the following steps. We'll be using this account to complete the deployment.
{% endhint %}

{% hint style="danger" %}
For production environments its best practice to create a specific 'installation user' account with the required role / permissions / privileges.
{% endhint %}

1. Run update & upgrade (optional).

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Add new user to system.

```bash
sudo adduser pentaho
```

3. Set password.

```
New password: password
Retype new password: password
passwd: password updated successfully
```

```
Follow the prompts to set the new user’s information. It is fine to accept the defaults to leave 
all this information blank:
```

4. Add new user to sudo group.

```bash
 sudo usermod -aG sudo pentaho
```

5. Test access.

```bash
su - pentaho
```

```bash
sudo ls -la /root
```

{% endtab %}

{% tab title="2. Open JDK & JRE 17" %}
{% hint style="info" %}

#### Java

Check the components reference ..!!

Pentaho 10.2 is certified on: Oracle OpenJDK & JRE 11 & 17
{% endhint %}

{% embed url="<https://docs.pentaho.com/pdia-10.2-install/components-reference>" %}

1. Run update & upgrade.

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Check whether Java is already installed in our system.

```bash
java -version
```

<figure><img src="/files/l1Rz6L1D16wPZ5DLhxB6" alt=""><figcaption><p>Java versions</p></figcaption></figure>

3. Install openjdk 11.

```bash
sudo apt install openjdk-17-jdk && sudo apt install openjdk-17-jre-headless
```

4. Check Java version.

```bash
java -version
```

```
openjdk 17.0.12 2024-07-16
OpenJDK Runtime Environment (build 17.0.12+7-Ubuntu-1ubuntu222.04)
OpenJDK 64-Bit Server VM (build 17.0.12+7-Ubuntu-1ubuntu222.04, mixed mode, sharing)
```

4. Tidy up.

```bash
sudo apt autoremove
```

***

{% hint style="info" %}

#### Set Java Version

You can have multiple Java installations on one server. You can configure which version is the default for use on the command line by using the update-alternatives command.
{% endhint %}

1. Run the following command to set the preferred Java version.

```bash
sudo update-alternatives --config java
```

```
There is only one alternative in link group java (providing /usr/bin/java): /usr/lib/jvm/java-17-openjdk-amd64/bin/java
Nothing to configure.
```

{% endtab %}

{% tab title="3. PENTAHO\_JAVA\_HOME" %}
{% hint style="info" %}

#### Set PENTAHO\_JAVA\_HOME

Perform the following steps to set the PENTAHO\_JAVA\_HOME environment variable. This will ensure that if there are multiple versions of Java the correct version is associated with PENTAHO\_JAVA\_HOME.
{% endhint %}

1. Run update & upgrade.

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Verify java version.

```bash
java -version
```

3. Determine path to OpenJDK.

```bash
update-alternatives --list java
```

***

{% hint style="info" %}

#### Set PENTAHO\_JAVA\_HOME - Global

As Pentaho is being installed under root to /opt/ then set PENTAHO\_JAVA\_HOME for all users.
{% endhint %}

1. Edit /etc/environment.

```bash
sudo nano /etc/environment
```

2. Add the following.

```bash
# set PENTAHO_JAVA_HOME
PENTAHO_JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64
```

3. Save.

```bash
CTRL + o
Enter
CTRL + x
```

4. Check path.

```bash
echo $PENTAHO_JAVA_HOME
```

5. Restart the server.

```
reboot
```

***

{% hint style="info" %}

#### Set PENTAHO\_JAVA\_HOME - Local

For just the specific user.
{% endhint %}

1. Set the path to OpenJDK.

```bash
sudo nano .bashrc
```

2. Add the following to the bottom.

```
# set PENTAHO_JAVA_HOME
export PENTAHO_JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64
```

3. Save.

```bash
CTRL + o
Enter
CTRL + x
```

4. Reload .bashrc.

```bash
. ~/.bashrc
```

5. Check path.

```bash
echo $PENTAHO_JAVA_HOME
```

6. Restart the server.

```
reboot
```

{% endtab %}

{% tab title="4. PostgreSQL 14 & 15" %}
{% hint style="info" %}

#### Pentaho Repository

The following databases are supported types as your Pentaho Repository 10.2.
{% endhint %}

<table><thead><tr><th width="261">Certified</th><th>Supported</th></tr></thead><tbody><tr><td>PostgreSQL 15</td><td>PostgreSQL 14 &#x26; 15</td></tr><tr><td>MySQL 8.026</td><td>MySQL 8.026</td></tr><tr><td>Oracle 23c</td><td>Oracle 19c &#x26; 23c (including patched versions)</td></tr><tr><td>MS SQL Server 2019</td><td>Microsoft SQL Server 2017 &#x26; 2019 (including patched versions)</td></tr><tr><td>Maria DB 11.1.2</td><td>Maria DB 11.1.2</td></tr></tbody></table>

{% tabs %}
{% tab title="4.1 Supported  Version" %}
{% hint style="info" %}

#### Supported Pentaho Database - PostgreSQL

Ensure you have a 'clean' PostgeSQL environment to avoid any potential conflicts - usually port.
{% endhint %}

1. Run update & upgrade.

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Check Postgresql version.

```bash
apt show postgresql -a
```

```
Package: postgresql
Version: 14+238
Priority: optional
Section: database
Source: postgresql-common (238)
Origin: Ubuntu
```

***

{% hint style="warning" %}
If PostgreSQL is installed, check its a supported version:

PDI v10.2 == Postgres 14 & 15
{% endhint %}

1. If not purge current Postgresql instance.

{% hint style="warning" %}
Execute each command separately ..

sudo apt autoremove to tidy up ..
{% endhint %}

```bash
sudo apt-get --purge remove postgresql
sudo apt-get purge postgresql*
sudo apt-get --purge remove postgresql postgresql-doc postgresql-common
```

2. Check for packages.

```bash
dpkg -l |grep postgres;PostgreSQL
```

{% endtab %}

{% tab title="4.2 Install PostgreSQL 14 & 15" %}
{% hint style="info" %}

#### Install PostgreSQL 14

PostgreSQL 14 is a default in the Unbuntu 22.04 apt repository.
{% endhint %}

1. Update Repository list.

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Install PostgreSQL 14.

```bash
sudo apt-get -y install postgresql-14 
```

```
Reading package lists... Done
Building dependency tree... Done
Reading state information... Done
The following additional packages will be installed:
  libcommon-sense-perl libjson-perl libjson-xs-perl libllvm14 libpq5
  libtypes-serialiser-perl postgresql-client-14 postgresql-client-common
  postgresql-common sysstat
Suggested packages:
  postgresql-doc-14 isag
The following NEW packages will be installed
  libcommon-sense-perl libjson-perl libjson-xs-perl libllvm14 libpq5
  libtypes-serialiser-perl postgresql-14 postgresql-client-14
  postgresql-client-common postgresql-common sysstat
0 to upgrade, 11 to newly install, 0 to remove and 0 not to upgrade.
Need to get 42.4 MB of archives.
...
```

3. Check the PostgreSQL version.

```bash
psql --version
```

```
psql (PostgreSQL) 14.12 (Ubuntu 14.12-0ubuntu0.22.04.1)
```

***

{% hint style="info" %}

#### PostgreSQL 15 - Add APT Repository

If you decide to go the PostgreSQL 15, the package is not available in the default package repository, so needs to be added to the official package repository using following commands.
{% endhint %}

1. Add APT Repository.

```bash
sudo sh -c 'echo "deb http://apt.postgresql.org/pub/repos/apt $(lsb_release -cs)-pgdg main" > /etc/apt/sources.list.d/pgdg.list'
```

2. Import GPG key.

```bash
wget -qO- https://www.postgresql.org/media/keys/ACCC4CF8.asc | sudo tee /etc/apt/trusted.gpg.d/pgdg.asc &>/dev/null
```

3. Update Repository list.

```bash
sudo apt update -y && sudo apt upgrade -y
```

4. Reboot server.

```bash
reboot
```

***

{% hint style="info" %}

#### Install PostgreSQL 15

Once package repository is updated, then ..
{% endhint %}

1. Install postgresql 15.

```bash
sudo apt install postgresql-client-15 postgresql-15
```

```
Reading package lists... Done
Building dependency tree... Done
Reading state information... Done
The following additional packages will be installed:
  libcommon-sense-perl libjson-perl libjson-xs-perl libpq5
  libtypes-serialiser-perl postgresql-client-common postgresql-common sysstat
Suggested packages:
  postgresql-doc-15 isag
The following NEW packages will be installed
  libcommon-sense-perl libjson-perl libjson-xs-perl libpq5
  libtypes-serialiser-perl postgresql-15 postgresql-client-15
  postgresql-client-common postgresql-common sysstat
0 to upgrade, 10 to newly install, 0 to remove and 0 not to upgrade.
Need to get 20.0 MB of archives.
After this operation, 66.6 MB of additional disk space will be used.
Do you want to continue? [Y/n] y
...
```

2. Check postgresql 15 status.

```bash
sudo systemctl status postgresql
```

```
pentaho@pentaho:~$ sudo systemctl status postgresql
● postgresql.service - PostgreSQL RDBMS
     Loaded: loaded (/lib/systemd/system/postgresql.service; enabled; vendor pr>
     Active: active (exited) since Fri 2024-08-16 00:38:15 BST; 2min 27s ago
   Main PID: 5687 (code=exited, status=0/SUCCESS)
        CPU: 1ms

Aug 16 00:38:15 pentaho systemd[1]: Starting PostgreSQL RDBMS...
Aug 16 00:38:15 pentaho systemd[1]: Finished PostgreSQL RDBMS.
```

```
sudo service postgresql stop // Stop the service
sudo service postgresql start // Start the service
sudo service postgresql restart // Stop and restart the service
sudo service postgresql reload // Reload the configuration without 
stopping the service
```

3. Check the PostgreSQL version.

```bash
psql --version
```

4. Tidy up installation.

```bash
sudo apt autoremove
```

{% endtab %}

{% tab title="4.3 Superuser" %}
{% hint style="info" %}

#### Grant privileges to a SuperUser

During installation, a 'postgres' user is created automatically. This user has full **superadmin** access to your entire PostgreSQL instance. Before you switch to this account, your logged in system user should have sudo privileges.
{% endhint %}

1. Log in as 'postgres' user.

```bash
sudo su -l postgres
```

2. Change password.

```plsql
psql -c "ALTER USER postgres WITH PASSWORD 'Welcome123'";
```

3. Exit

```
CTRL + d
```

***

{% hint style="info" %}

#### Pentaho Superuser

Grant access & privileges to Pentaho superuser.
{% endhint %}

1. Switch to 'postgres' user.

```bash
sudo -u postgres psql
```

2. Create 'pentaho' user.

```bash
CREATE USER pentaho WITH PASSWORD 'Welcome123';
```

3. Upgrade 'pentaho' to superuser.

```bash
ALTER USER pentaho WITH SUPERUSER;
```

4. Exit.

```
CTRL + d
```

5. Switch to 'pentaho' user:

```bash
 sudo -u pentaho psql postgres
```

6. Check connection details.

```bash
\conninfo
```

```
postgres=# \conninfo
You are connected to database "postgres" as user "pentaho" via socket in "/var/run/postgresql" at port "5432".
```

7. Exit.

```bash
CTRL + d
```

***

{% hint style="info" %}

#### Allow Remote Connections

<mark style="color:green;">For reference only as connecting to Postgresql via localhost</mark>

By default, PostgreSQL accepts connections from the localhost only. However, we can easily modify the configuration to allow connection from remote clients.

PostgreSQL reads its configuration from the postgresql.conf file which is located: /etc/postgresql/\<version>/main/ directory
{% endhint %}

1. Edit the postgresql.conf file.

```bash
cd
cd /etc/postgresql/15/main/
sudo nano postgresql.conf
```

2. Uncomment the line that starts with the listen\_addresses, and replace ‘localhost’ with ‘\*’.

{% hint style="info" %}
Alternatively, you can specify a specific IP address or a range of IP addresses that are allowed to connect to the server.
{% endhint %}

3. Modify the pg\_hba.conf file.

```bash
sudo nano /etc/postgresql/15/main/pg_hba.conf
```

4. Locate the following section and modify

```
# IPv4 local connections: 
host    all             all             127.0.0.1/32         md5 
```

```
# IPv4 local connections:
host    all             all             0.0.0.0/0            md5 
```

5. Allow port 5432 through the firewall.

```bash
sudo ufw allow 5432/tcp
```

6. Restart Postgres.

```bash
sudo service postgresql restart
```

{% endtab %}

{% tab title="4.4 pgAdmin" %}
{% hint style="info" %}

#### pgAdmin4 Desktop

The **pgadmin4** package contains both:

• **pgadmin4-desktop** – Provides desktop application for Ubuntu system.

• **pgadmin4-web** – Provides the web interface accessible in a web browser
{% endhint %}

1. Update Repository list.

```bash
sudo apt update -y && sudo apt upgrade -y
```

2. Install Public Key.

```bash
curl -fsS https://www.pgadmin.org/static/packages_pgadmin_org.pub | sudo gpg --dearmor -o /usr/share/keyrings/packages-pgadmin-org.gpg
```

3. Create Repository config file.

```bash
sudo sh -c 'echo "deb [signed-by=/usr/share/keyrings/packages-pgadmin-org.gpg] https://ftp.postgresql.org/pub/pgadmin/pgadmin4/apt/$(lsb_release -cs) pgadmin4 main" > /etc/apt/sources.list.d/pgadmin4.list && apt update'
```

4. Install pgAdmin4 for both desktop - see below for Web version.

```bash
sudo apt install pgadmin4-desktop
```

5. Check contents APT Repository.

```bash
cat /etc/apt/sources.list.d/pgadmin4.list
```

***

{% hint style="info" %}

#### Pentaho Server Group

In pgAdmin, a server group is a way to organize and categorize your PostgreSQL server connections. It's essentially a logical grouping mechanism that helps you manage multiple database server connections more efficiently.
{% endhint %}

1. Click on: 'Add New Server'.

<figure><img src="/files/vKPwcEgMN4uVDjPsdBvP" alt=""><figcaption><p>create server group</p></figcaption></figure>

2. Enter: Pentaho for server name.

<figure><img src="/files/L0xfEBZnf8zm7JzXtXNo" alt=""><figcaption><p>server group</p></figcaption></figure>

3. Click on 'Connection'.

<figure><img src="/files/6YybXscGxg5OPTXTIjpY" alt=""><figcaption><p>Connection details</p></figcaption></figure>

```
Welcome123
```

{% hint style="danger" %}
Don't save the password in Production environments.
{% endhint %}

4. Save.

<figure><img src="/files/c9HuvcEZni3GT9xj9Qvd" alt=""><figcaption><p>pgAdmin4 UI</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Install Pentaho Server

Installation of Pentaho Server components ..

{% hint style="info" %}

#### Pentaho Server Installation

This section will guide you through the installation of the Pentaho Server:

* create installation directories
* create Pentaho Repository databases
* configure JDBC database connections
* start Pentaho server - systemd
* license manager
  {% endhint %}

<figure><img src="/files/uGcROs2ByszHNvuGQu2y" alt=""><figcaption><p>Pentaho Pro Suite</p></figcaption></figure>

{% tabs %}
{% tab title="1. Pentaho Server" %}
{% hint style="info" %}

#### Pentaho Server Directories

The Pentaho server is a web application that runs in an Apache Tomcat servlet container.
{% endhint %}

1. Create /opt/pentaho server directories.

```bash
cd
sudo mkdir -p /opt/pentaho/{server,software}
```

```
* server        - server zip packages
* software      - Pentaho binaries
```

2. Create /opt/pentaho/software sub-directories.

```bash
cd
cd /opt/pentaho/software
sudo mkdir -p {server,shims,ee-plugins,db_drivers}
```

```
* server       - server binaries
* shims        - collections of Hadoop libraries required to communicate with a specific version of Hadoop
* ee-pligins   - pentaho ee-plugins
* db_drivers   - database drivers
```

***

{% hint style="info" %}

#### Unpack Pentaho Server Packages

The jar command is a general-purpose archiving and compression tool, based on ZIP and the ZLIB compression format.

x - Extract files from a JAR archive

f - Sets the file specified by the jarfile operand to be the name of the JAR file that is created
{% endhint %}

1. Copy Pentaho server package.

```bash
cd
cd ~/Downloads/'Archive Build (Suggested Installation Method)'/
sudo cp * /opt/pentaho/software/server
```

2. Unjar pentaho-server-ee-10.2.0.0-222.zip to /opt/pentaho/server/

```bash
cd
cd /opt/pentaho/server
sudo jar -vxf /opt/pentaho/software/server/pentaho-server-ee-10.2.0.0-222.zip
```

3. Change the permission for all .sh files.

{% hint style="info" %}
All .sh files will need executable permission.
{% endhint %}

```bash
cd
cd /opt/pentaho/server
sudo find . -iname "*.sh" -exec bash -c 'chmod +x "$0"' {} \;
```

```
find [obvious!]
.        - from this folder. You can put a path instead
-iname   - case insensitive name
"*.sh"   - wildcard filename
-exec    - utility to execute commands
bash     - what tool you want to use (you can use sh instead)
-c flag means execute the following command as interpreted by this program.
chmod +x - command to change the file to executable
"$0"     - The value that was passed to the utility
{}       - If the string {} appears anywhere in the utility name or the arguments it is replaced by the pathname of the current file.
;        - Terminates the command
```

3. Check that it matches the following directory structure:

{% hint style="info" %}
/opt/pentaho/

server/

pentaho-server/

pentaho-solutions/

system

The server plugins are installed into the system folder.
{% endhint %}
{% endtab %}

{% tab title="2. Pentaho Repository" %}
{% hint style="warning" %}

#### Pentaho Repository

The Pentaho Repository resides on the database that you installed during the Windows or Linux environment preparation step, and consists of the following components:

#### **Jackrabbit**

Contains the solution repository, examples, security data, and content data from reports that you use Pentaho software to create.

#### **Quartz**

Holds data that is related to scheduling reports and jobs.

#### **Hibernate**

Holds data that is related to audit logging.

#### **Pentaho Operations Mart**

Report on system usage and performance.
{% endhint %}

{% tabs %}
{% tab title="2.1 Database Passwords" %}
{% hint style="warning" %}
For your production server, Pentaho recommends that you change the default passwords in the following SQL script files to make the databases more secure.
{% endhint %}

1. Examine postgresql database scripts.

```bash
cd
cd /opt/pentaho/server/pentaho-server/data
ls -l
cd postgresql
ls -l
```

{% hint style="info" %}
For this workshop we're going to keep the defaults user and password.
{% endhint %}

```
create_jcr_postgresql.sql
create_quartz_postgresql.sql
create_repository_postgresql.sql
pentaho_logging_postgresql.sql
pentaho_mart_postgresql.sql
```

2. Examine postgresql database scripts.

```bash
 cd
 cd /opt/pentaho/server/pentaho-server/data/postgresql
 cat create_jcr_postgresql.sql
```

```
--
-- note: this script assumes pg_hba.conf is configured correctly
--

-- \connect postgres postgres

drop database if exists jackrabbit;
drop user if exists jcr_user;

CREATE USER jcr_user PASSWORD 'password';
CREATE DATABASE jackrabbit WITH OWNER = jcr_user ENCODING = 'UTF8' TABLESPACE =>
GRANT ALL PRIVILEGES ON DATABASE jackrabbit to jcr_user;
```

3. Exit.

```
CTRL x
```

***

{% hint style="danger" %}

#### Change Authentication Mode

Client authentication is controlled by pg\_hba.conf and is stored in the database cluster's data directory. (HBA stands for host-based authentication.)

You need ensure that users defined in the scripts are able to be authenticated to connect to the tables.
{% endhint %}

1. Edit pg\_hba.conf.

```bash
cd
sudo nano /etc/postgresql/15/main/pg_hba.conf
```

2. Replace 'peer' for local connection with 'md5'.

<figure><img src="/files/SgVJzJciJZrdQNmyZMSp" alt=""><figcaption><p>Change from peer to md5</p></figcaption></figure>

{% hint style="info" %}
The MD5 (message-digest algorithm) hashing algorithm is a one-way cryptographic function that accepts a message of any length as input and returns as output a fixed-length digest value to be used for authenticating the original message.
{% endhint %}

3. Save.

```
CTRL + o
Enter
CTRL + x
```

4. Restart service.

```bash
sudo service postgresql restart
```

{% endtab %}

{% tab title="2.2 SQL Scripts" %}
{% hint style="info" %}

#### Run SQL Scripts

You will find SQL scripts for each supported Repository database:

* PostgreSQL
* Oracle
* MS Sql Server
* MySQL
* MariaDB
  {% endhint %}

1. Check postgresql repository is running.

```bash
sudo systemctl status postgresql
```

2. Select postgresql scripts directory.

```bash
cd
cd /opt/pentaho/server/pentaho-server/data/postgresql
ls -l
```

```
-rw-r--r-- 1 root root   464 Aug  7 07:42 alter_script_postgresql_BISERVER-13674.sql
-rw-r--r-- 1 root root   363 Aug  7 07:40 create_jcr_postgresql.sql
-rw-r--r-- 1 root root  5647 Aug  7 07:40 create_quartz_postgresql.sql
-rw-r--r-- 1 root root   356 Aug  7 07:40 create_repository_postgresql.sql
-rw-r--r-- 1 root root  4016 Aug  7 07:43 pentaho_logging_postgresql.sql
-rw-r--r-- 1 root root  1220 Aug  7 07:43 pentaho_mart_drop_postgresql.sql
-rw-r--r-- 1 root root 19035 Aug  7 07:43 pentaho_mart_postgresql.sql
-rw-r--r-- 1 root root   286 Aug  7 07:43 pentaho_mart_upgrade_audit_postgresql.sql
-rw-r--r-- 1 root root  7533 Aug  7 07:43 pentaho_mart_upgrade_postgresql.sql
```

3. Log in as 'pentaho' superuser.

```bash
sudo -su pentaho psql postgres
```

```
Passw0rd123
```

💡The default password for each user: password

💡You can switch database in PostgreSQL with the command: \c

<table><thead><tr><th width="310">Script</th><th width="118">Database</th><th>User / Password</th></tr></thead><tbody><tr><td>\i create_jcr_postgresql.sql</td><td>Jackrabbit</td><td></td></tr><tr><td>\i create_quartz_postgresql.sql</td><td>Quartz</td><td>pentaho_user/password</td></tr><tr><td>\i create_repository_postgresql.sql</td><td>Hibernate</td><td></td></tr><tr><td>\i pentaho_mart_postgresql.sql</td><td>OpsMart</td><td>hibuser/password</td></tr><tr><td>\i pentaho_logging_postgresql.sql</td><td>Logging</td><td>hibuser/password</td></tr></tbody></table>

{% hint style="warning" %}
Ensure you're in the

/opt/pentaho/server/pentaho-server/data/postgresql directory.

Refer to the table above for username / password.

You may need to \q and log back in.
{% endhint %}

4. Execute the scripts.

```
\i create_jcr_postgresql.sql
```

```
\i create_quartz_postgresql.sql
```

Enter password:

```
password
```

Quit:

```
\q
```

Log back in as 'pentaho' superuser.

```bash
sudo -su pentaho psql postgres
```

```
Welcome123
```

```
\i create_repository_postgresql.sql
```

```
\i pentaho_mart_postgresql.sql
```

Enter password:

```
password
```

Quit:

```
\q
```

Log back in as 'pentaho' superuser.

```bash
sudo -su pentaho psql postgres
```

```
Welcome123
```

```
\i pentaho_logging_postgresql.sql
```

Enter password:

```
 password
```

Quit:

```
\q
```

4. Check the databases in pgAdmin.

{% hint style="warning" %}
No tables are created in the Hibernate and Jackrabbit databases. These are created during the installation of the Pentaho server.
{% endhint %}

<figure><img src="/files/lLN20Zq8UrIlzWid2NwZ" alt=""><figcaption><p>Pentaho Repository databases</p></figcaption></figure>
{% endtab %}

{% tab title="2.3 Configure Repository" %}
{% hint style="info" %}

#### Pentaho Repository

Now that you have initialized your repository database, you will need to configure Quartz, Hibernate, Jackrabbit, and Pentaho Operations Mart for a PostgreSQL database.

PostgreSQL is configured by default; if you kept the default passwords and port, you will not need to set up Quartz, Hibernate, Jackrabbit or the Pentaho Operations Mart.
{% endhint %}

{% hint style="danger" %}
By default, the examples in this section are for a PostgreSQL database that runs on port 5432. The default password is also in these examples.

If you have a different port or different password, make sure that you change the password and port number in these examples to match the ones in your configuration.
{% endhint %}

***

{% hint style="info" %}

#### Quartz

Event information, such as scheduled reports, is stored in the Quartz JobStore. During the installation process, you must indicate where the JobStore is located by modifying the quartz.properties file.
{% endhint %}

1. Navigate to quartz directory.

```bash
cd 
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system/scheduler-plugin/quartz
sudo nano -c quartz.properties
```

2. Locate the #\_replace\_jobstore\_properties section and check.

```
[line 300/451]
org.quartz.jobStore.driverDelegateClass =
org.quartz.impl.jdbcjobstore.PostgreSQLDelegate
```

3. Locate the # Configure Datasources section and check.

```
[line 379/451]
org.quartz.dataSource.myDS.jndiURL = Quartz
```

4. Exit.

```
CTRL + x
```

***

{% hint style="info" %}

#### Hibernate

Modify the Hibernate settings file to specify where Pentaho should find the Pentaho Repository’s Hibernate configuration file. The Hibernate configuration file specifies driver and connection information, as well as dialects and how to handle connection closes and timeouts.

The Hibernate database is also where the Pentaho Server stores the audit logs that act as source data for the Pentaho Operations Mart.
{% endhint %}

1. Navigate to hibernate-settings directory.

```bash
cd 
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system/hibernate
sudo nano -c hibernate-settings.xml
```

2. Locate the config-file section and check.

```
[line 36/53]
<config-file>system/hibernate/postgresql.hibernate.cfg.xml</config-file>
```

3. Exit.

```bash
CTRL + x
```

4. Display postgresql.hibernate.cfg.xml.

```bash
sudo nano -c postgresql.hibernate.cfg.xml
```

5. Check postgresql has been set as default.

```
[line 35/51]
    <!--  Postgres 8 Configuration -->
    <property name="connection.driver_class">org.postgresql.Driver</property>
    <property name="dialect">org.hibernate.dialect.PostgreSQLDialect</property>
    <property name="hibernate.connection.datasource">java:comp/env/jdbc/Hibernat>
    <property name="connection.pool_size">10</property>
    <property name="show_sql">false</property>
    <property name="hibernate.jdbc.use_streams_for_binary">true</property>
    <!-- replaces DefinitionVersionManager -->
    <property name="hibernate.hbm2ddl.auto">update</property>
    <!-- load resource from classpath -->
    <mapping resource="hibernate/postgresql.hbm.xml" />
    <!-- mapping resource above is from CE; below is from EE -->
    <mapping resource="hibernate/postgresql.EE.hbm.xml" />
  </session-factory>
</hibernate-configuration>

```

6. Exit.

```
CTRL + x
```

***

{% hint style="info" %}

#### Jackrabbit

Apache Jackrabbit is a platform of java open source content repository. A JCR (Java content repository) is a type of object database to customizing, storing, searching and retrieving hierarchical data.
{% endhint %}

{% embed url="<https://jackrabbit.apache.org/jcr/index.html>" %}

{% hint style="info" %}
As shown in the table below, locate and verify or change the code so that the PostgreSQL lines are not commented out, but the MySQL, Oracle, and MS SQL Server lines are commented out.

If you have a different port or different password, make sure that you change the password and port number in these examples to match the ones in your configuration.
{% endhint %}

1. Navigate to jackrabbit directory.

```bash
cd 
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system/jackrabbit
sudo nano -c repository.xml
```

<table><thead><tr><th width="151">Line #</th><th width="227">Section</th><th>Check</th></tr></thead><tbody><tr><td>line 71/442</td><td>Repository</td><td>Filesystem schema: postgresql</td></tr><tr><td>line 129/442</td><td>Datastore</td><td>databaseType: postgresql</td></tr><tr><td>line 231/442</td><td>Workspace</td><td>Filesystem schema: postgresql</td></tr><tr><td>line 279/442</td><td>Persistence Manager (1)</td><td>PersistenceManager schema: postgesql</td></tr><tr><td>line 347/442</td><td>Versioning</td><td>Filesystem schema: postgresql</td></tr><tr><td>line 398/442</td><td>Persistence Manager (2)</td><td>PersistenceManager schema: postgesql</td></tr><tr><td>line 434/442</td><td>Database Journal</td><td>Journal schema: postgresql</td></tr></tbody></table>

2. Exit.

```
CTRL + x
```

{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="3. Tomcat" %}
{% hint style="info" %}

#### Tomcat

After your Repository has been configured, you must configure the web application servers to connect to the Pentaho Repository. In this step, you will make JDBC and JNDI connections to the Hibernate, Jackrabbit, and Quartz components.
{% endhint %}

{% tabs %}
{% tab title="3.1 Database Drivers" %}
{% hint style="warning" %}

#### Database Drivers

To connect to a database, including the Pentaho Repository database, you will need to download and install a JDBC driver to the appropriate places for Pentaho components as well as on the the web application server that contains the Pentaho Server.

Due to licensing restrictions, Pentaho cannot redistribute some third-party database drivers. You must download the file yourself and install it yourself.
{% endhint %}

{% hint style="info" %}
For this workshop we're going to distribute a MySQL driver ..
{% endhint %}

{% embed url="<https://docs.pentaho.com/pdia-10.2-install/jdbc-drivers-reference>" %}

1. Copy the JDBC drivers to jdbc-distribution.

```bash
cd
cd ~/Downloads/'Database Drivers'/
sudo cp * /opt/pentaho/software/db_drivers
```

```bash
cd 
cd /opt/pentaho/software/db_drivers
sudo cp mysql-connector-j-9.0.0.jar /opt/pentaho/server/jdbc-distribution
```

2. Distribute the drivers.

```bash
 cd
 cd /opt/pentaho/server/jdbc-distribution
 sudo ./distribute-files.sh /opt/pentaho/server/pentaho-server/tomcat/lib
```

```
/opt/pentaho/server/jdbc-distribution
DEBUG: Using PENTAHO_JAVA_HOME
DEBUG: _PENTAHO_JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64
DEBUG: _PENTAHO_JAVA=/usr/lib/jvm/java-17-openjdk-amd64/bin/java
You must restart your Pentaho Server and Client tools to begin using the new drivers.
```

{% hint style="danger" %}
You will need to restart the Pentaho Server and Client Tools to register the driver.

Multiple distribution paths can be set, separated by a 'space'.
{% endhint %}

```bash
reboot
```

Location of JDBC drivers in Pentaho+:

| Server / Design Tool               | Directory                                  |
| ---------------------------------- | ------------------------------------------ |
| Pentaho Server                     | /server/pentaho-server/tomcat/lib          |
| Pentaho Data Integration (Spoon)   | /design-tools/data-integration/lib         |
| Pentaho Report Designer (PRD)      | /design-tools/report-designer/lib/jdbc     |
| Pentaho Aggregation Designer (PAD) | /design-tools/aggregation-designer/drivers |
| Pentaho Schema Workbench (PSW)     | /design-tools/schema-workbench/drivers     |
| Pentaho Metadata Editor (PME)      | /design-tools/metadata-editor/libext/JDBC  |

{% hint style="warning" %}
Check that the driver(s) have been added .. sometimes the distribution tool fails.. !!
{% endhint %}
{% endtab %}

{% tab title="3.2 context.xml" %}
{% hint style="info" %}

#### context.xml

Database connection and network information, such as the username, password, driver class information, IP address or domain name, and port numbers for your Pentaho Repository database are stored in the context.xml file.
{% endhint %}

1. Check context.xml.

```bash
cd
cd /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/META-INF
sudo nano -c context.xml
```

{% hint style="warning" %}
In a Production environment, check the username, password, driver class information, IP address (or domain name), and port numbers to match the correct values for your environment.
{% endhint %}

2. Exit.

```
CTRL + x
```

{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="4. Start Server" %}
{% hint style="info" %}

#### Start Server

Now that you have completed the initial Pentaho Archive installation steps, you are ready to start the Pentaho Server.
{% endhint %}

1. Switch to pentaho-server directory.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./start-pentaho.sh
```

```
DEBUG: Using PENTAHO_JAVA_HOME
DEBUG: _PENTAHO_JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64
DEBUG: _PENTAHO_JAVA=/usr/lib/jvm/java-17-openjdk-amd64/bin/java
DEBUG: PENTAHO_LICENSE_INFORMATION_PATH=
Using CATALINA_BASE:   /opt/pentaho/server/pentaho-server/tomcat
Using CATALINA_HOME:   /opt/pentaho/server/pentaho-server/tomcat
Using CATALINA_TMPDIR: /opt/pentaho/server/pentaho-server/tomcat/temp
Using JRE_HOME:        /usr
Using CLASSPATH:       /opt/pentaho/server/pentaho-server/tomcat/bin/bootstrap.jar:/opt/pentaho/server/pentaho-server/tomcat/bin/tomcat-juli.jar
Using CATALINA_OPTS:   -Xms2048m -Xmx6144m -Djava.library.path=/opt/pentaho/server/pentaho-server/pentaho-solutions/native-lib/linux/x86_64/ -Dsun.rmi.dgc.client.gcInterval=3600000 -Dsun.rmi.dgc.server.gcInterval=3600000 -Dfile.encoding=utf8 -Djava.locale.providers=COMPAT,SPI -DDI_HOME="/opt/pentaho/server/pentaho-server/pentaho-solutions/system/kettle"
Tomcat started.
```

3. Tail the log (new terminal).

```bash
cd
sudo tail -f /opt/pentaho/server/pentaho-server/tomcat/logs/catalina.2024-*.log
```

```
....
31-Jul-2024 17:20:03.810 INFO [main] org.apache.coyote.AbstractProtocol.start Starting ProtocolHandler ["http-nio-8080"]
31-Jul-2024 17:20:03.827 INFO [main] org.apache.catalina.startup.Catalina.start Server startup in [61207] milliseconds
```

***

{% hint style="info" %}

#### Pentaho User Console

The Pentaho User Console (PUC) is a web-based design environment where you can analyze data, create interactive reports, dashboard reports, and build integrated dashboards to share business intelligence solutions with others in your organization and on the internet.
{% endhint %}

{% embed url="<http://localhost:8080/pentaho>" %}
Link to Pentaho Server
{% endembed %}

```
Username: admin
Password: password
```

<figure><img src="/files/w77BwfNFY8r44hcl7CgT" alt=""><figcaption><p>Pentaho User Console</p></figcaption></figure>

***

{% hint style="info" %}

#### Systemd

Systemd is a system and service manager for Linux operating systems. It's designed to be backwards compatible with SysV init scripts, and provides several features to start system services in parallel, which can reduce boot times.
{% endhint %}

To use this service:

1. Save the file below as: `/etc/systemd/system/pentaho-server.service`

```systemd
[Unit]										  
Description=Pentaho Server							  
Before=multi-user.target															 								  
Before=graphical.target								  
After=network.service								  
After=network.target								  
After=syslog.target			  
											  
[Service]								       	  
Type=forking									  
Environment="JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64"		    					  
ExecStart=/opt/pentaho/server/pentaho-server/start-pentaho.sh
ExecStartPost=/bin/echo pentaho...end of unitfile		 		  
ExecStop=/opt/pentaho/server/pentaho-server/stop-pentaho.sh
TimeoutSec=500									  
IgnoreSIGPIPE=no									  
KillMode=process									  
GuessMainPID=no									  
RemainAfterExit=yes								  
SuccessExitStatus=5 6								  
User=root									  
 										  
[Install]									  
WantedBy=multi-user.target
```

2. Reload the systemd daemon.

```bash
sudo systemctl daemon-reload
```

3. Start the service.

```bash
sudo systemctl start pentaho-server
```

4. Enable the service to start on boot.

```bash
sudo systemctl enable pentaho-server
```

```
Created symlink /etc/systemd/system/multi-user.target.wants/pentaho-server.service → /etc/systemd/system/pentaho-server.service.
```

You can then manage the service using standard systemd commands:

To stop:

```bash
sudo systemctl stop pentaho-server
```

To restart:

```bash
sudo systemctl restart pentaho-server
```

To check status:

```bash
sudo systemctl status pentaho-server
```

{% endtab %}

{% tab title="5. License Manager" %}

<figure><img src="/files/mOWgyPRr3bAV3ReZu4UB" alt=""><figcaption><p>Licensing</p></figcaption></figure>

{% tabs %}
{% tab title="License Manager" %}
{% hint style="warning" %}
The Pentaho Licensing model has changed in Pentaho Pro Suite 10.+ Licenses are now handled via a License Manager.

This embedded service (on-prem / cloud) will enable our customers (Direct & OEM) to manage their PDI & BA entitlements with greater visibility and ease.

The License manger also checks EE plugins.
{% endhint %}

{% hint style="info" %}

#### Trial license

You can obtain a Pentaho trial license to test the product before you acquire it. You need internet access to activate a trial license and run the installed trial version. It is not possible to run it for an extended period of time while disconnected from internet.
{% endhint %}

The temporary license expires thirty days after the start of the evaluation period.

1. To request a trial activation ID, go to the [Pentaho Enterprise download page](https://www.hitachivantara.com/en-us/products/pentaho-platform/data-integration-analytics/download-pentaho.html). Alternatively, contact the Pentaho Sales team.
2. On the download page, click Start a Free 30 Day Trial to open the trial registration form.
3. Complete the trial registration form and click Submit. The trial entitlement and activation ID is sent to you by email.
4. Download and install the Pentaho product. When launching the product, the Add License window opens.

<figure><img src="/files/hkq4eQpHcLVUvkHv7y0x" alt=""><figcaption><p>License Manager</p></figcaption></figure>

5. Select Activation Code.
6. Copy the activation ID that you received from the local license manager, paste it into the provided field, and click OK

<figure><img src="/files/MIo0Oyt29pANx7uWLsyy" alt=""><figcaption><p>Trial license - activation code</p></figcaption></figure>

{% hint style="warning" %}

#### Enterprise licenses

If you are an existing customer wanting to upgrade from Pentaho 9.x or earlier supported versions, do not start the server before upgrading the licenses. You must install the new version of the product before activating the licenses.
{% endhint %}

{% hint style="info" %}
**Activate a license using a cloud license server**

If you are able to access our cloud license server without any security restrictions, this is the quickest way to get up and running with Pentaho.

Copy the cloud license server URL that Pentaho emails you into the License Server field of the Add License dialog box that opens when you launch the product and click OK.
{% endhint %}
{% endtab %}

{% tab title="Set ENV License Path" %}
{% hint style="info" %}

#### Set ENV License Path

To ensure that the Pentaho Server uses the same location to store and retrieve your Pentaho licenses, you must create a PENTAHO\_LICENSE\_INFORMATION\_PATH system environment variable for your Pentaho user account if it does not exist. It does not matter what location you choose; however, the location needs to be available to the user account(s) that run the Pentaho Server.
{% endhint %}

Perform the following steps to set the environment variable for the license path in Linux.

1. Edit the /etc/environment file.

```bash
cd
cd /etc
sudo nano environment
```

2. Add this line in a convenient place (changing the path if necessary)

```bash
# 
export PENTAHO_LICENSE_INFORMATION_PATH=/home/pentaho/.pentaho/.elmLicInfo.plt
```

{% hint style="info" %}
The license information file is saved in the /home/pentaho/.pentaho folder.
{% endhint %}

3. Log out and log back into the operating system for the change to take effect.
4. Verify that the variable is properly set using the following command.

```bash
env | grep PENTAHO_LICENSE_INFORMATION_PATH
```

The PENTAHO\_LICENSE\_INFORMATION\_PATH variable is now set.
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Server Plugins

Installation of Reporting Plugins ..

{% hint style="info" %}

#### Pentaho Server Reporting Plugins

The Pentaho Analyzer (PAZ) plugin installation involves downloading the appropriate plugin files from the Pentaho marketplace or official repository and extracting them to the server's plugin directory, typically located in the pentaho-server/pentaho-solutions/system folder. After copying the files, you'll need to restart the Pentaho Server to register the new plugin. The Analyzer plugin provides OLAP analysis capabilities, allowing users to create interactive pivot tables and charts from multidimensional data sources.

The Pentaho Interactive Reports (PIR) plugin follows a similar installation process, where the plugin files are downloaded and placed in the designated plugin directory structure. This plugin extends Pentaho's reporting capabilities by enabling users to create and modify reports directly within the web interface, providing drag-and-drop functionality for report design and real-time data visualization without requiring the separate Report Designer tool.

The Pentaho Dashboard Designer (PDD) plugin installation requires extracting the dashboard designer files to the appropriate system plugin folder and ensuring proper permissions are set. This plugin integrates seamlessly with the Pentaho User Console, providing a web-based interface for creating interactive dashboards that can incorporate various data visualizations, reports, and analysis components. After installation, users can access the dashboard designer through the PUC interface to build and deploy custom dashboards for their business intelligence needs.
{% endhint %}

{% hint style="info" %}

#### Unpack Pentaho Server Plugin Packages

The jar command is a general-purpose archiving and compression tool, based on ZIP and the ZLIB compression format.

x - Extract files from a JAR archive

f - Sets the file specified by the jarfile operand to be the name of the JAR file that is created
{% endhint %}

{% tabs %}
{% tab title="1. Analyzer" %}
{% hint style="info" %}

#### Analyzer Reports

Analyzer Reports is an intuitive analytical visualization tool that filters and drills down into business information contained in Pentaho analysis data sources. Use Analyzer Reports if you want to compile data quickly in an interactive environment, perform advanced sorting and filtering of your data, and want to see chart visualizations that include conditional stop-lighting.
{% endhint %}

1. Stop Pentaho Server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Check for server plugins.

```bash
cd
cd /opt/pentaho/software/server
ls
```

3. Unjar paz-plugin-ee-10.2.0.0-222.zip.

```bash
cd
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system
sudo jar -vxf /opt/pentaho/software/server/paz-plugin-ee-10.2.0.0-222.zip
```

4. Check that it matches the following directory structure:

{% hint style="info" %}
/opt/pentaho

server

pentaho-server

pentaho-solutions

system

analyzer
{% endhint %}

5. Start Pentaho server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}

<figure><img src="/files/ALhngzLMq5biBAxBpQ2s" alt=""><figcaption><p>Analyzer report</p></figcaption></figure>
{% endtab %}

{% tab title="2. Interactive Reports" %}
{% hint style="info" %}

#### Interactive Reports

Pentaho Interactive Reporting is a drag-and-drop, browser-based design environment for interactive reports that allows you to quickly add elements to your report and format them to your preference.
{% endhint %}

1. Stop Pentaho Server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Check for server plugins.

```bash
cd
cd /opt/pentaho/software/server
ls
```

3. Unjar pir-plugin-ee-10.2.0.0-222.zip.

```bash
cd
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system
sudo jar -vxf /opt/pentaho/software/server/pir-plugin-ee-10.2.0.0-222.zip
```

4. Check that it matches the following directory structure:

{% hint style="info" %}
\~/Pentaho

server

pentaho-server

pentaho-solutions

system

analyzer

pentaho-interactive-reporting
{% endhint %}

5. Start Pentaho server.

```bash
cd
cd /opt/entaho/server/pentaho-server
sudo ./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}

<figure><img src="/files/vRBTPyfojdwtsr1wALCF" alt=""><figcaption><p>Interactive report</p></figcaption></figure>
{% endtab %}

{% tab title="3. Dashboard Designer" %}
{% hint style="info" %}

#### Dashboard Designer

Creating a dashboard in Dashboard Designer is as simple as choosing a layout template, theme, and the content you want to display. In addition to displaying content generated from Interactive Reports and Analyzer, Dashboard Designer can also include these content types.

* **Charts**: simple bar, line, area, pie, and dial charts created with Chart Designer
* **Data Tables**: tabular data
* **URLs**: Web sites that you want to display in a dashboard panel

Dashboard Designer has dynamic filter controls, which enable dashboard viewers to change a dashboard's details by choosing different values from a drop-down list, and to control the content in one dashboard panel by changing the options in another. This is known as content linking.
{% endhint %}

1. Stop Pentaho Server.

```bash
cd
cd /opt/entaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Check for server plugins.

```bash
cd
cd /opt/entaho/software/server
ls
```

3. Unjar pdd-plugin-ee-10.2.0.0-222.zip.

```bash
cd
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system
sudo jar -vxf /opt/pentaho/software/server/pdd-plugin-ee-10.2.0.0-222.zip
```

4. Check that it matches the following directory structure:

{% hint style="info" %}
/opt/pentaho

server

pentaho-server

pentaho-solutions

system

analyzer

pentaho-interactive-reporting

dashboards
{% endhint %}

5. Start pentaho server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./start-pentaho.sh
```

{% embed url="<http://localhost:8080/pentaho>" %}

<figure><img src="/files/G7tiGxIb33CtcGnqFxDG" alt=""><figcaption><p>Dashboard Designer</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Install Client Tools

Installation of Clients Tools ..

{% hint style="info" %}

#### Pentaho Client Tools

There are two methods for installing the Business Analytics (BA) design tools. You can use either of these methods:

* Pentaho Business Analytics Evaluation Wizard - Windows Desktop.
* Install each separate tool manually - Linux / Windows Desktop.

The Pentaho Business Analytics Installation Wizard is the easiest way to install design tools, utilities, or plugins on the server or client workstations. The manual method allows you to manually copy design tool installation files to any directory on the server or client workstations. The deployment pattern will depend on the DevOps requirements.
{% endhint %}

<figure><img src="/files/uGcROs2ByszHNvuGQu2y" alt=""><figcaption><p>Pentaho Pro Suite</p></figcaption></figure>

The following steps install the client plugins into a Linux Desktop - for [**Windows**](/pentaho-10-installation/installation/evaluation-installation)

{% hint style="info" %}

#### Unpack Client Package

The jar command is a general-purpose archiving and compression tool, based on ZIP and the ZLIB compression format.

x - Extract files from a JAR archive

f - Sets the file specified by the jarfile operand to be the name of the JAR file that is created
{% endhint %}

{% tabs %}
{% tab title="1. Data Integration" %}
{% hint style="info" %}

#### Pentaho Data Integration

Pentaho Data Integration (PDI) provides the Extract, Transform, and Load (ETL) capabilities that facilitates the process of capturing, cleansing, and storing data using a uniform and consistent format that is accessible and relevant to end users and IoT technologies.
{% endhint %}

#### Pentaho Client Directories

1. Create \~/Pentaho/design-tools directory.

```bash
cd
mkdir -p ~/Pentaho/design-tools
```

2. Check for client plugins.

```bash
cd
cd ~/Downloads/'Client Tools'/'PDI (Spoon)'
ls
```

3. Unjar pdi-ee-client-10.2.0.0-222.zip.

```bash
cd
cd ~/Pentaho/design-tools
jar -vxf ~/Downloads/'Client Tools'/'PDI (Spoon)'/pdi-ee-client-10.2.0.0-222.zip
```

4. Change the permission for all .sh files.

{% hint style="info" %}
All .sh files will need executable permission.
{% endhint %}

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec bash -c 'chmod +x "$0"' {} \;
```

5. Check that it matches the following directory structure:

{% hint style="info" %}
\~/Pentaho

design-tools

data-integration

jdbc-distribution

license-installer

PDI plugins are installed in the plugins folder.
{% endhint %}

***

{% hint style="info" %}

#### Data Integration UI

Ubuntu 22.04 the Repository is:

• missing libwebgtk: webkit browser extensions.

• and fails to load canberra-gtk-module
{% endhint %}

1. Add package repository.

```bash
sudo apt-get install -qq software-properties-common
```

2. Add repository entry.

```bash
sudo apt-key adv --keyserver keyserver.ubuntu.com --recv-keys 3B4FE6ACC0B21F32
sudo add-apt-repository 'deb [trusted=yes] http://cz.archive.ubuntu.com/ubuntu bionic main universe'
```

3. Update repositories.

```bash
sudo apt-get update
```

5. Install package.

```bash
sudo apt-get install -qq libwebkitgtk-1.0-0
```

```bash
sudo apt-get install libcanberra-gtk-module
```

6. Start PDI.

```bash
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

<figure><img src="/files/lSY5V6EdHNTzT6AmD7n7" alt=""><figcaption><p>Spoon UI</p></figcaption></figure>
{% endtab %}

{% tab title="2. Metadata Editor" %}
{% hint style="info" %}

#### Pentaho Metadata Editor

Pentaho Metadata Editor is a tool that allows you to create and manage metadata models and domains for Pentaho data integration and analytics. Metadata models and domains are logical representations of your physical data sources that make them easier to understand and use by business users.
{% endhint %}

1. Unjar pme-ee-10.2.0.0-222.zip.

```bash
cd
cd ~/Pentaho/design-tools
jar -vxf ~/Downloads/'Client Tools'/'Metadata Editor'/pme-ee-10.2.0.0-222.zip
```

2. Change the permission for all .sh files.

{% hint style="info" %}
All .sh files will need executable permission.
{% endhint %}

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec bash -c 'chmod +x "$0"' {} \;
```

3. Check that it matches the following directory structure:

{% hint style="info" %}
\~/Pentaho

design-tools

data-integration

jdbc-distribution

license-installer

metadata-editor

The PME plugins are installed in the plugins folder.
{% endhint %}

***

{% hint style="info" %}

#### Metadata Editor

Building a Pentaho Metadata Editor model starts with connecting to your data sources and defining business tables that provide user-friendly views of your underlying database tables. You rename technical columns to business terms, create relationships between tables, and can add calculated fields and hierarchies that make sense to end users.

The model building process includes setting up security rules to control data access and adding localization support for multiple languages. You can hide technical complexity while exposing only the relevant data and metrics that business users need for their analysis and reporting.

Once complete, the metadata model is published to the Pentaho server where it serves as a semantic layer for reporting tools, dashboards, and ad-hoc analysis. This provides business users with a consistent, governed view of organizational data without requiring them to understand the underlying database structure or write complex queries.
{% endhint %}

1. Start PME.

```bash
cd
cd ~/Pentaho/design-tools/metadata-editor
./metadata-editor.sh
```

<figure><img src="/files/tfDLaoXeDghFew8nCp0Y" alt=""><figcaption><p>Pentaho Metadata Editor</p></figcaption></figure>
{% endtab %}

{% tab title="3. Schema WorkBench" %}
{% hint style="info" %}

#### Schema Workbench

Pentaho Schema Workbench is a tool that allows you to create and test Mondrian OLAP cube schemas visually. Mondrian OLAP cube schemas are XML files that define the logical structure of a multidimensional database, such as dimensions, hierarchies, levels, measures, and calculations.
{% endhint %}

1. Unjar psw-ee-10.2.0.0-222.zip.

```bash
cd
cd ~/Pentaho/design-tools
jar -vxf ~/Downloads/'Client Tools'/'Schema Workbench'/psw-ee-10.2.0.0-222.zip
```

2. Change the permission for all .sh files.

{% hint style="info" %}
All .sh files will need executable permission.
{% endhint %}

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec bash -c 'chmod +x "$0"' {} \;
```

2. Check that it matches the following directory structure:

{% hint style="info" %}
\~/Pentaho

design-tools

data-integration

jdbc-distribution

license-installer

metadata-editor

schema-workbench

The PSW plugins are installed in the plugins folder.
{% endhint %}

***

{% hint style="info" %}

#### Schema Workbench

Building a Pentaho Schema Workbench schema starts with connecting to your data warehouse and defining the basic schema structure. You create cubes that represent your main analytical subjects, each containing dimensions for context (like time or geography) and measures for quantitative data (like sales or revenue).

The development process involves mapping these logical structures to your database tables by defining fact tables for measures and dimension tables for attributes. You establish proper join relationships between tables and create hierarchies within dimensions to enable drill-down analysis from high-level summaries to detailed data.

After configuring the core structure, you can add calculated members and advanced features to enhance analytical capabilities. The final step involves validating the schema to ensure all relationships work correctly, then deploying it to the Pentaho server where reporting and analysis tools can access the multidimensional data model.
{% endhint %}

1. Start PSW.

```bash
cd
cd ~/Pentaho/design-tools/schema-workbench
./workbench.sh
```

<figure><img src="/files/xm0svKRIIA7KbC4eLMTQ" alt=""><figcaption><p>Schema Workbench</p></figcaption></figure>
{% endtab %}

{% tab title="4. Aggregation Designer" %}
{% hint style="info" %}

#### Aggregation Designer

Pentaho Aggregation Designer is a tool that helps you to improve the performance of your Pentaho Analyzer (Mondrian) OLAP cubes by creating and deploying aggregate tables. Aggregate tables are tables that contain pre-aggregated measures from the base fact table, which can speed up the query processing by reducing the amount of data that needs to be scanned and aggregated.
{% endhint %}

1. Unjar pad-ee-10.2.0.0-222.zip.

```bash
cd
cd ~/Pentaho/design-tools
jar -vxf ~/Downloads/'Client Tools'/'Aggregation Designer'/pad-ee-10.2.0.0-222.zip
```

2. Change the permission for all .sh files.

{% hint style="info" %}
All .sh files will need executable permission.
{% endhint %}

```bash
cd
cd ~/Pentaho/design-tools
find . -iname "*.sh" -exec bash -c 'chmod +x "$0"' {} \;
```

3. Check that it matches the following directory structure:

{% hint style="info" %}
\~/Pentaho

design-tools

data-integration

jdbc-distribution

license-installer

metadata-editor

schema-workbench

aggregation-designer

The PAD plugins are installed in the plugins folder.
{% endhint %}

***

{% hint style="info" %}

#### Aggregation Designer

The Pentaho Aggregation Designer is a performance optimization tool that creates pre-calculated summary tables (aggregates) for OLAP cubes to speed up analytical queries. It automatically analyzes query patterns and recommends which aggregate tables to build based on frequently accessed data combinations.

The tool connects to existing OLAP schemas, examines dimensional hierarchies and measures, then generates the physical aggregate tables in the database along with updated schema definitions. This significantly reduces query response times by allowing the system to serve results from pre-computed summaries rather than calculating them on-demand from large fact tables.

Building it requires components for schema analysis, performance modeling, aggregate recommendation algorithms, and deployment utilities that handle the database changes and schema updates needed to implement the optimizations.
{% endhint %}

1. Start PAD.

```bash
cd
cd ~/Pentaho/design-tools/aggregation-designer
./startaggregationdesigner.sh
```

<figure><img src="/files/EleCwcSnqNOfVuspHh51" alt=""><figcaption><p>Aggregation Designer</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Post Installation Tasks

Hardening & performance ..

{% hint style="info" %}
The Pentaho Server has options that must be set manually, outside of the Administration page of the User Console.
{% endhint %}

<details>

<summary>Remove server banner</summary>

Most web servers display its version and modules in use by default. Best security practices recommend that you disable this option, since it can be used to find vulnerabilities of your site.

1. Edit: \<tomcat installed directory>/conf/server.xml file.

```
cd
cd /opt/pentaho/server/
sudo nano server.xml
```

2. Add following under Connector port and save the file

```
Server =” “
<Connector port="8080" protocol="HTTP/1.1"
connectionTimeout="20000"
Server =" "
redirectPort="8443" />
```

3.

x

</details>

<details>

<summary>Starting Tomcat with a Security Manager</summary>

Security Manager protects you from an untrusted applet running in your browser, use of a SecurityManager, while running Tomcat can protect your server from trojan servlets, JSPs, JSP beans, and tag libraries or even inadvertent mistakes.

```
start tomcat with –security argument
<tomcat installed directory>/bin# ./startup.sh -security
```

</details>

<details>

<summary>Change the Web Application Name</summary>

These instructions only work on Tomcat servers that are configured to accept context.xml overrides built into deployed .war files.

1. Stop the Pentaho Server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Edit context.xml.

```bash
cd
cd /opt/pentaho/server/pentaho-server/tomcat/webapps/pentaho/META-INF
sudo nano context.xml
```

3. Change the context.

```
<context path="/company" docbase="webapps/company/">
```

4. Save.

```
CTRL + O
Enter
CTRL + x
```

5. Change directory name to the same context name:

\~/Pentaho/server/pentaho-server/tomcat/webapps

In this example, rename the pentaho folder to *company*.

6. Edit the main .jsp page.

```bash
cd
cd /opt/pentaho/server/pentaho-server/tomcat/webapps/ROOT
sudo nano index.jsp
```

```
....
html
  <head>
    <title>Pentaho Business Analytics</title>
    <meta http-equiv="refresh" content="0;URL=/company">
  </head>
  ....
```

7. Finally change: fully-qualified-server-url.

```bash
cd
cd /opt/pentaho/pentaho-server/pentaho-solutions/system
sudo nano server.properties
```

```
....
# FullyQualifiedServerUrl is used only in the case of offline content generation
# and whenever something need to talk back to the server
fully-qualified-server-url=http://localhost:8080/company/
....
```

</details>

<details>

<summary>Change Port Number</summary>

Pentaho server default port is 8080.

1. Stop the Pentaho Server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sudo ./stop-pentaho.sh
```

2. Navigate to.

```bash
cd
cd /opt/pentaho/server/pentaho-server/tomcat/conf/
sudo nano server.xml
```

3. Change the port number in the connector from 8080.

```
    <Connector URIEncoding="UTF-8"
       port="8090" protocol="HTTP/1.1"
       connectionTimeout="20000"
       redirectPort="8443"
       relaxedPathChars="[]|"
       relaxedQueryChars="^{}[]|&amp;"
       maxHttpHeaderSize="65536"
    />
```

4. Save.

```bash
Ctrl + o
Enter
Ctrl + x
```

5. Navigate to.

```bash
cd
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system 
sudo nano server.properties
```

6. Change the port number to match the port number set in the connector.

```
fully-qualified-server-url=http://localhost:8090/pentaho/
```

7. Save.

```bash
Ctrl + o
Enter
Ctrl + x
```

8. Restart Pentaho server.

```bash
cd
cd /opt/pentaho/server/pentaho-server
sh stop-pentaho.sh 
```

x

</details>

<details>

<summary>Change SHUTDOWN port and Command</summary>

By default, tomcat is configured to be shutdown on 8005 port. Do you know you can shutdown tomcat instance by doing a telnet to IP:port and issuing SHUTDOWN command?

```
# telnet localhost 8005
Trying ::1... telnet:
connect to address ::1:
Connection refused Trying 127.0.0.1...
Connected to localhost.
Escape character is '^]'.
SHUTDOWN Connection closed by foreign host.
#
```

You see having default configuration leads to high-security risk. It’s recommended to change tomcat shutdown port and default command to something unpredictable.

1. Edit Go to $tomcat/conf/server.xml file.

ii. Modify server.xml by using vim editor

```
<Server port="8005" shutdown="SHUTDOWN">
```

The default shutdown port and command must be changed or it should be disabled.

</details>

<details>

<summary>Replace default 404, 403, 500 page</summary>

Having default page for not found, forbidden, server error exposes Tomcat version and that leads to security risk if you are running with vulnerable version. Let’s look at default 404 page.

To mitigate, you can first create a general error page and configure web.xml to redirect to general error page.

1. Go to $tomcat/webapps/$application
2. Create an error.jsp file

```html
<html>
<head>
<title>404-Page Not Found</title>
</head>
<body> That's an error! </body>
</html>
```

3. Go to $tomcat/conf folder

* Add following in web.xml by using vi. Ensure you add before \</web-app> syntax

```
<error-page>
<error-code>404</error-code>
<location>/error.jsp</location>
</error-page>
<error-page>
<error-code>403</error-code>
<location>/error.jsp</location>
</error-page>
<error-page>
<error-code>500</error-code>
<location>/error.jsp</location>
</error-page>
```

Restart tomcat server. Now, let’s test it.

</details>

<details>

<summary>Session Timeout</summary>

The session timeout for all web applications must be set to 20 minutes.\
This can be done by editing the file in the $tomcat/conf/web.xml and setting the following configuration option:

```
<session-config>
<session-timeout>20</session-timeout>
</session-config>
```

</details>

<details>

<summary>Change the Karaf Startup Timeout Settin<strong>g</strong></summary>

Upon start up, the system waits for Karaf to install all of its features before timing out. If you modify Karaf and it now takes longer to install during start up, you may need to extend the default timeout setting to allow Karaf more time to install. The current default timeout is 2 minutes (120000 milliseconds).

You can change this default timeout by editing the server.properties file.

1\. Stop the Pentaho Server.

2\. Navigate to the /pentaho-server/pentaho-solutions/system directory.

3\. Open the server.properties file with any text editor, and search for the karafWaitForBoot parameter.

4\. Uncomment the line containing the parameter and set it to your desired wait time in milliseconds.

\# This sets the amount of time the system will wait for karaf to install all of\
\# it’s features before timing out. The default value is 2 minutes but can be\
\# overridden here.\
\#karafWaitForBoot = 120000

5\. Save and close the file.

6\. Restart the Pentaho Server.

</details>

<details>

<summary>Remove Sample Data from the Pentaho Server</summary>

By default, Pentaho provides a sample data source and a solution directory filled with example content. These samples are provided for evaluation and testing. Once you are ready to move from an evaluation or testing scenario to development or production, you can remove the sample content.

Follow the instructions below to completely remove the Pentaho sample data and solutions:

1. Stop the Pentaho Server.

```bash
sudo systemctl stop pentaho-server
```

ii. Delete the samples.zip file from the /pentaho-server/pentaho-solutions/system/default-content directory. If you performed a manual WAR build and deployment, then the file path is /pentaho-server/pentaho-solutions/system.

iii. Edit the /pentaho/WEB-INF/web.xml file inside of the deployed pentaho.war. As laid down by the Pentaho Plus graphical installer and archive packages, this path should be /pentaho-server/tomcat/webapps/pentaho/WEB-INF/web.xml. If you performed a manual WAR build and deployment, then you must adjust the path to fit your configuration.

iv. Remove the hsqldb-databases section from the /pentaho/WEB-INF/web.xml file:

v. BEGIN HSQLDB DATABASES

```
<!-- [BEGIN HSQLDB DATABASES] -->
<context-param>
<param-name>hsqldb-databases</param-name>
<param-value>sampledata@../../data/hsqldb/sampledata</param-value>
</context-param>
<!-- [END HSQLDB DATABASES] -->
```

vi. Remove the hsqldb-starter section from the /pentaho/WEB-INF/web.xml file:

vii. BEGIN HSQLDB STARTER

```
<!-- [BEGIN HSQLDB STARTER] -->
<listener>
<listener-class>org.pentaho.platform.web.http.context.HsqldbStartupListener</listener-class>
</listener>
<!-- [END HSQLDB STARTER] -->
```

viii. Remove the SystemStatusFilter:

**Note:** This is not part of the Pentaho samples; it provides error status messages that are only useful for development and testing purposes, and should be removed from a production system.

```
<filter>
<filter-name>SystemStatusFilter</filter-name>
<filter-class>com.pentaho.ui.servlet.SystemStatusFilter</filter-class>
<init-param>
<param-name>initFailurePage</param-name>
<param-value>InitFailure</param-value>
<description>This page is displayed if the Pentaho+ System fails to properly initialize.</description>
</init-param>
</filter>
```

i. Save and close the web.xml file.

ii. Delete the /pentaho-server/data/ directory. This directory does not exist if you installed Pentaho with the installation wizard. It contains a sample database, control scripts for that database, the environment settings it needs to run, and SQL scripts to initialize a new repository.

iii. Restart the Pentaho+ Server.

iv. Log on to the User Console with the administrator user name and password and go to the Browse Files page.

1. In the Folders pane, expand the Public folder and click to highlight the folder containing the Steel Wheels sample data. Click Move to Trash in the Folder Actions pane and confirm the deletion.
2. Highlight the folder containing the Pentaho Plus Operations Mart sample data. Click Move to Trash in the Folder Actions pane and confirm the deletion.

Your Pentaho+ Server instance is now cleaned of samples and development/testing pieces, and is streamlined for production.

</details>

<details>

<summary>Disable Home Perspective Widgets</summary>

The User Console default Home perspective contains the Getting Started widget, which has easy instructions and tutorials for evaluators.

Perform the following steps to hide not only the Getting Started widget, but also other Home perspective widgets.

1. Shut down the Pentaho Server if it is currently running.
   1. Choose one of the following options depending on your deployment status:

***

* If you have not yet deployed, navigate to:

/pentaho-platform/user-console/source/org/pentaho/mantle/home/properties/config.properties file.

***

* If you have manually deployed and want to hide widgets later, navigate to:

/pentaho-server/tomcat/webapps/pentaho/mantle/home/properties/config.properties file.

3. Find the line that starts with disabled-widgets= and type in the ID of the widget getting-started, as shown in the following example:

```
disabled-widgets=getting-started,recents,favorites
```

4\. Save and close the file.

***

You can also hide the Recents and Favorites widgets using the same method.

1. Locate the /pentaho-server/tomcat/webapps/pentaho/mantle/home directory and open the index.jsp file with any text editor.
2. Find the following line of code and comment it out, then save and close the file.

```javascript
<script language='JavaScript' type='text/javascript' src='http://admin.brightcove.com/js/BrightcoveExperiences.js'></script>
```

3. Start the Pentaho Server and log in to the User Console.

</details>

<details>

<summary>Turn Autocomplete Off for Web App Login Screen</summary>

The User Console's sign-in settings have autocomplete turned on by default.

Perform the following steps to manually turn off the autocompletion functionality:

1. Stop the Pentaho Server.
2. Modify PUCLogin.jsp.

```bash
cd
sudo nano /opt/pentaho/server/pentaho-server/tomact/webapps/pentaho/jsp/PUCLogin.jsp
```

3. Locate the following sections of code and change the autocomplete entry to off, as shown:

```
<input id="j_username" name="j_username" type="text" placeholder="" autocomplete="off">
```

```
<input id="j_password" name="j_password" type="password" placeholder="" autocomplete="off">
```

3. Save and close the PUCLogin.jsp file.
4. Restart the Pentaho Server.

</details>

<details>

<summary>Set System Max Row Limit for Interactive Reports</summary>

You can prevent too many resources from hitting your database server at once by setting a system-wide maximum row-limit for Pentaho Interactive Reports. Your users can still define their own design-time row limits in PIR, but they will never be able to go over the maximum number of rows that you have specified while designing their reports.

1. Stop the Pentaho Server.

```bash
sudo systemctl stop pentaho-server
```

2. Edit: /opt/pentaho/server/pentaho-server/pentaho-solutions/system/pentaho-interactive-reporting/settings.xml file.

```bash
cd
cd /opt/pentaho/server/pentaho-server/pentaho-solutions/system/pentaho-interactive-reporting
sudo nano settings.xml
```

3. Navigate to: \<query-limit> tag and change the default number of 100000 within the tags to the maximum number of rows desired.

\<!– The maximum number of rows that will be rendered in a report on PIR edit and view mode. A zero value means no limit. –>

```
<query-limit>100000</query-limit>
```

4. Save.

```
CTRL + O
Enter
CTRL + X
```

5. Start the Pentaho Server.

```bash
sudo systemctl start pentaho-server
```

If you are migrating content from a previous version, you will need to add the \<query-limit> tag to your settings.xml for PIR.

</details>

<details>

<summary>Increase the CSV File Upload Limit</summary>

You may find that you need to increase the size of the upload limit for your CSV files. These steps guide you through this process.

1. Go to /pentaho-server/pentaho-solutions/system and open the pentaho.xml file.

Edit the XML as needed (sizes are measured in bytes):

```
<file-upload-defaults>
<relative-path>/system/metadata/csvfiles/</relative-path>
<!-- max-file-limit is the maximum file size, in bytes, to allow to be uploaded to the server -->
<max-file-limit>10000000</max-file-limit>
<!-- max-folder-limit is the maximum combined size of all files in the upload folder, in bytes. -->
<max-folder-limit>500000000</max-folder-limit>
</file-upload-defaults>
```

2. Save your changes to the file.
3. In the User Console, go to Tools > Refresh System Settings to ensure that the change is implemented.
4. Restart the User Console.

**Change the Staging Database for CSV Files**

Hibernate is the default staging database for CSV files. Follow these instructions if you want to change the staging database.

1. Go to /pentaho-solutions/system/data-access and open the settings.xml file with any text editor.
2. Edit the settings.xml file as needed. The default value is shown in the sample below.

iii. \<!– settings for Agile Data Access –>

iv. \<data-access-staging-jndi>hibernate\</data-access-staging-jndi>

This value can be a JNDI name or the name of a Pentaho Database Connection.

3. Save and close the file.
4. Restart the User Console

</details>


# Pentaho Upgrade & Patches

Pentaho Upgrade Installer ..

{% hint style="info" %}

#### Pentaho Upgrade Installer

You can upgrade your Pentaho products from version 8.3 or later to version 10.1 with the Pentaho Upgrade Installer.

The upgrade installer checks your environment for version 8.3 or later Pentaho products, creates a backup of these products, then upgrades them to version 10.1. The Pentaho Upgrade Installer works for any Pentaho products you have installed on your server or workstations, including your Pentaho Server and your Pentaho client tools.

The Pentaho Upgrade Installer requires 22 GB of free space to perform the upgrade process.
{% endhint %}

{% hint style="warning" %}
If you need to upgrade your Pentaho products from a version earlier than 8.3, such as 7.1 or higher, you must upgrade your products to version 8.3, then use the Pentaho Upgrade Installer to move from version 8.3 to 10.1.
{% endhint %}

{% tabs %}
{% tab title="Checklist" %}
{% hint style="info" %}
Before you can run the Pentaho Upgrade Installer, you must also perform the following tasks:
{% endhint %}

* [x] Verify that your system components are current.

{% embed url="<https://docs.pentaho.com/pdia-10.2-install/components-reference>" %}

* [x] If you are upgrading an environment that includes the Pentaho Server, stop the server prior to performing backups and installation.

```bash
cd 
cd ~/[Pentaho Installation Directory]/server/pentaho-server
sh stop-pentaho.sh
```

1. Review your customizations. During the upgrade process, you can help the upgrade installer specify which items contain your customizations. See [Specify customized items to address after upgrading](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/r2F5~x211KC0wwU0qcq6cQ) for details. Then, after upgrading your Pentaho products to 10.1, you can merge your previous customizations into post-upgrade versions of the Pentaho files. See the [Apply customizations](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/S2d88cUzhPmuc8jUpi9NaA) post-upgrade task for instructions.

*

{% hint style="warning" %}
The upgrade process does not retain the drivers for your Hadoop clusters. You will need to re-install your drivers after completing the upgrade process.
{% endhint %}

1. Note: The upgrade process does not retain the drivers for your Hadoop clusters. You will need to re-install your drivers after completing the upgrade process. See the [Install drivers for your Hadoop clusters](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/UEFngjwGGT~SZRKXjCWeNQ) post-upgrade task for details.
2. If you are using plugins with your Pentaho products, review and back up your plugins to a separate directory structure.

{% hint style="warning" %}
The upgrade process does not retain your plugins. You will need to re-apply your plugins after completing the upgrade process.
{% endhint %}

1.
2. See the [Apply your plugins](https://docs.hitachivantara.com/r/HuHAFx8OjcQg31CW~6gISg/pKp_UIrYWFitwgefW2h9ig) post-upgrade task for details.
3. If you are upgrading the Pentaho Server, verify that no users are logged on to the server.As a best practice, perform the upgrade process of the Pentaho Server during off-business hours to minimize the impact on your day-to-day operations.
4. Before installing the Pentaho Upgrade, verify that you have the most recent version of Java installed and that the JAVA\_HOME environment variable is set to that version of Java.
   {% endtab %}

{% tab title="Release" %}
{% hint style="info" %}
You can upgrade your Pentaho products from version 8.3 or later to version 9.4 using the Pentaho Upgrade Installer.

The upgrade installer checks your environment for version 8.3 or later Pentaho products, creates a backup of these products, then upgrades them to version 9.4.

The Pentaho Upgrade Installer works for any Pentaho products you have installed on your server or workstations, including your Pentaho Server and your Pentaho client tools.
{% endhint %}

{% hint style="warning" %}
The Pentaho Upgrade Installer requires 22 GB of free space to perform the upgrade process.
{% endhint %}

Before you can run the Pentaho Upgrade Installer, you must also perform the following tasks:

* [ ] Verify that your system components are current.

{% embed url="<https://help.hitachivantara.com/Documentation/Pentaho/9.4/Setup/Components_Reference>" %}
Pentaho 9.4
{% endembed %}

* [ ] If you are upgrading an environment that includes the Pentaho Server, stop the server prior to performing backups and installation.

```bash
cd 
cd ~/Pentaho/server/pentaho-server
sh stop-pentaho.sh
```

* [ ] Review any customizations.

During the upgrade process, you can help the upgrade installer specify which items contain your customizations.
{% endtab %}
{% endtabs %}


# Windows Installation

Windows Installation ..

{% hint style="info" %}
The Installation Wizard provides the easiest and quickest way to install Pentaho Pro Suite on Windows.

In this workshop, you will install 'everything' enabling access to a 'local' Pentaho Repository.
{% endhint %}

#### Windows Installation

1. Open Windows Explorer and navigate to the downloaded pentaho-business-analytics-10.2.0-222.exe installation file.
2. Double-click the pentaho-business-analytics-10.2.0-222.exe file to launch it.
3. Ignore the Antivirus warning ..

<figure><img src="/files/JwHnv8INAt35CemI0y2N" alt="" width="563"><figcaption></figcaption></figure>

4. Accept the License agreement.

<figure><img src="/files/4UWfPphzVLEbE9pNxPJE" alt="" width="563"><figcaption><p>End user license agreement</p></figcaption></figure>

5. Keep the default installation path.

<figure><img src="/files/c4rRr03H5J9YE33WhgC4" alt="" width="563"><figcaption><p>Default installation path</p></figcaption></figure>

6. The default Pentaho Repository database is PostgreSQL. Enter the password.

<figure><img src="/files/zvdNATTIBTzuSnUENjmC" alt="" width="563"><figcaption><p>Pentaho Repository</p></figcaption></figure>

7. Decide if you wish to install everything or specific components.

<figure><img src="/files/SXPWx4oc7BeyOzyYH4Yu" alt="" width="563"><figcaption><p>Keep it simple.</p></figcaption></figure>

{% hint style="info" %}
With the Pentaho Installation Wizard you can choose one of two ways to install Pentaho components:

* Default: Select the Keep it simple. Give me everything option in the installation wizard.
* Custom: Select the Let me decide for myself option in the installation wizard.
  {% endhint %}

<figure><img src="/files/xUw6UJnuE4qedX2rpfrl" alt="" width="563"><figcaption><p>Let me decide for myself</p></figcaption></figure>

8. Enter either an Activation ID or a Licensing server URL.

<figure><img src="/files/kk5Y6o5Bfx57joyJosqQ" alt="" width="563"><figcaption></figcaption></figure>

9. Click 'Next' to start the installation.

<figure><img src="/files/bYuGGkY3z2g05R0rgJ6v" alt="" width="563"><figcaption></figcaption></figure>

{% hint style="warning" %}
Various notifications will appear during the installation process. Allow Pentaho services to be installed.
{% endhint %}

x

x

{% hint style="info" %}

```
Pentaho/
   /design-tools/
       /aggregation-designer/
       /data-integration/
       /metadata-editor/
       /report-designer/
       /schema-workbench/
   /documentation/
   /java/
   /jdbc-distribution/
   /license-installer/

pentaho/postgresql/
pentaho/scripts/
pentaho/server/
```

{% endhint %}

x

x

x


# EE Plugins

{% tabs %}
{% tab title="Customer Portal" %}

### Pentaho Customer Portal

{% hint style="info" %}
The Pentaho Customer Portal is a comprehensive support hub that provides Pentaho customers with technical resources, documentation, and support services for their data integration and analytics solutions. The portal offers access to product releases, service packs, support documentation, and world-class support services, along with best practices libraries covering topics like installation, security configuration, cloud deployment, and performance tuning. It supports both Normal Release and Long-Term Support (LTS) versions with Active Patching and Limited Support phases, serving as the central resource for customers using Pentaho's platform to access, prepare, and analyze data across various environments.

You will require Pentaho account details.
{% endhint %}

1. Log into Pentaho Customer Portal.

{% embed url="<https://support.pentaho.com/hc/en-us>" %}
Link to Pentaho Customer Portal
{% endembed %}

<figure><img src="/files/V2fdwIp9CR6r8UtirJgp" alt=""><figcaption><p>Pentaho Customer Portal</p></figcaption></figure>

2. Click on: Pentaho -> DOWNLOAD.
3. Click on: Pentaho 10.2 EE Marketplace Plugins Release.

<figure><img src="/files/9FW5gdDIdQiawpoBhu3E" alt=""><figcaption><p>Pentaho Downloads</p></figcaption></figure>

4. Select the required Pentaho plugin version to download.

<figure><img src="/files/4oHwPT4NeRHR65RVqslu" alt=""><figcaption><p>EE Pentaho Plugins - versions</p></figcaption></figure>

5. Select the plugins to download.

<figure><img src="/files/islkt8eIZcBh8uBe0aH4" alt=""><figcaption><p>EE Pentaho 10.2 Plugins</p></figcaption></figure>
{% endtab %}

{% tab title="EE Plugins" %}
x

x

x

{% tabs %}
{% tab title="Hierachical Data Type (HDT)" %}
{% hint style="info" %}
HDT is a new data type, provided as a plugin, that allows you to process hierarchical structure data in PDI and enables the ability to convert between HDT fields and formatted strings (JSON).
{% endhint %}

x

x

x
{% endtab %}

{% tab title="Second Tab" %}
x
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Ubuntu Pentaho Lab

Setup Pentaho Server + Plugins on Ubuntu ..

x

x

{% tabs %}
{% tab title="Ubuntu" %}
{% hint style="info" %}

#### Ubuntu 22.04 LTS

{% endhint %}

x

x

x
{% endtab %}

{% tab title="Download Packages" %}
x

x

x

{% tabs %}
{% tab title="Pentaho EE" %}
x

x

{% hint style="info" %}

#### Customer Portal

You can access the Pentaho packages with your principal account credentials.
{% endhint %}

1. Log into Pentaho Customer Portal.

{% embed url="<https://support.pentaho.com/hc/en-us>" %}
Pentaho Customer Portal
{% endembed %}

<figure><img src="/files/V2fdwIp9CR6r8UtirJgp" alt=""><figcaption><p>Customer Support Portal</p></figcaption></figure>

2. Click on Pentaho > DOWNLOAD option.
3. Select Pentaho version.

<figure><img src="/files/9FW5gdDIdQiawpoBhu3E" alt=""><figcaption><p>Pentaho Downloads</p></figcaption></figure>

4. Select the required Pentaho packages to download.

<figure><img src="/files/XKRJoD2ogmF5k43nIWkw" alt=""><figcaption><p>Packages</p></figcaption></figure>

{% hint style="info" %}

#### Pentaho Packages

Below is a list of the main packages used in the deployment.

**Server Components**

* pentaho-server-ee-10.2.0.0-222-dist.zip
* paz-plugin-ee-10.2.0.0-222-dist.zip
* pir-plugin-ee-10.2.0.0-222-dist.zip
* pdd-plugin-ee-10.2.0.0-222-dist.zip

**Client Components**

* pdi-ee-client-10.2.0.0-222-dist.zip
* psw-ee-client-10.2.0.0-222-dist.zip
* pme-ee-client-10.2.0.0-222-dist.zip
* prd-ee-client-10.2.0.0-222-dist.zip

**Utilities & Tools**

* dock-maker-10.2.0.0-222-dist.zip

**EE Plugins**

•
{% endhint %}

x

x
{% endtab %}

{% tab title="Pentaho CE" %}
x

x

x

x

x
{% endtab %}
{% endtabs %}

x

x
{% endtab %}
{% endtabs %}


# Common Questions

#### General

<details>

<summary>Where can I get a copy of the 'workshop files' ?</summary>

All the collateral can be found at: \~/Workshop--Installation.

You can also copy/fork the Git repository:

```
https://github.com/jporeilly/Workshop--Installation-Linux
```

```git
gh repo clone jporeilly/Workshop--Installation
```

</details>

<details>

<summary>I cant write to sampledata database?</summary>

For security reasons the privileges from V10 have been removed. You will need to install Docker Desktop and create a MySQL sampledata database container.

[Broken mention](broken://pages/nyuNuY53XaIkQsM4MkZX)

</details>

<details>

<summary>Why do we need Docker &#x26; Docker Compose?</summary>

Its easier to manage some of the required applications - databases, Brokers, etc .. - as Docker containers.

</details>

#### Windows - Linux - MacOS (Self-paced Labs)

<details>

<summary>Where can i download 30-day Pentaho Enterprise Edition?</summary>

An 30-day activation code will be emailed to you:

<https://pentaho.com/download/#download-pentaho>

For BYOL:

<https://pentaho.com/pentaho-ee-onprem/>

</details>

#### Linux (Instructor-led Labs)

<details>

<summary>VM is unresponsive !</summary>

* Refresh the browser session to reconnect.
* Try another browser. The recommended browser: Google Chrome.
* If you're connecting via a Corporate VPN, then this may cause issues. Contact your IT dept to get the URL 'white' listed.

</details>

<details>

<summary>When does the Lab expire ?</summary>

The initial duration is 5 days. You will receive an email asking if you wish to extend your time limit.

</details>

<details>

<summary>Videos aren't loading?</summary>

* Hard refresh your browser. CTRL + F5

</details>

<details>

<summary>Is there sound ?</summary>

Yes .. There's no sound card attched to the Lab, so you'll need to copy and paste the Lab Guide URL in your host machine browser. 😊

</details>


# Pentaho Support

How to work with Pentaho Support and open effective tickets ..

{% hint style="info" %}
**Overview**

Pentaho Support at Pentaho helps you troubleshoot product issues and answer usage questions. Use this page to understand who can open tickets, what to include, and how to contact Support.
{% endhint %}

{% embed url="<https://www.youtube.com/watch?v=RRXw1d09RMk>" %}
Support Onboarding
{% endembed %}

<details>

<summary>Pentaho Virtualization Support Statement</summary>

Pentaho supports virtualization technology across the Pentaho Suite. This statement explains what is supported and what may be requested during troubleshooting.

#### What is supported

Pentaho supports Pentaho products that run on:

* Supported operating systems
* Minimum hardware requirements

This applies whether you run Pentaho in a virtual environment or not.

#### What you are responsible for

You are responsible for failures caused by the hardware layer or operating system layer. This includes misuse of virtualization software.

#### What Support may ask you to do

Pentaho does not require you to reproduce every issue on physical hardware.

Pentaho may ask you to reproduce or diagnose an issue on a native (non-virtual) supported operating system.

This request is made only when there is reason to believe virtualization contributes to the issue.

#### How virtualization-related issues are handled

When a problem may relate to virtualization, investigation is handled as follows:

* Pentaho provides standard support for all products.
* If an issue occurs in a virtual environment, you may be asked to reproduce it on a physical (non-virtual) server. Pentaho then provides regular support.
* You can authorize Pentaho to investigate virtualization-related items at normal time and materials rates. If the issue is virtualization-related, you may request a software change, if a resolution is possible.
* If the issue is not virtualization-related, investigation and resolution are covered under regular maintenance.

Pentaho is expected to work in virtual environments.

#### Performance notes

Performance impacts may still occur. These impacts may or may not be caused by virtualization. They may fall outside this support statement.

</details>

<details>

<summary>Pentaho’s Global Data Protection &#x26; Privacy Policy</summary>

During troubleshooting, you may need to share data from your systems. This can include diagnostics, metadata, result sets, and similar artifacts.

#### Ways to send data

Pentaho Support provides several ways to transmit data, including:

* Email
* The Customer Portal
* Pentaho Content Platform (HCP)

#### Before you share

Share only what Support requests. Remove or redact sensitive data when possible.

See [Hitachi Vantara’s Global Data Protection & Privacy Policy](https://www.hitachivantara.com/en-us/legal) for details on how personal data and information are protected while in our possession.

</details>

{% tabs %}
{% tab title="Named Support Contacts" %}
{% hint style="info" %}
Pentaho Support works directly with your organization’s **Named Support Contacts**. These are specific people with a **unique (non-generic) email address** and a **phone number**.
{% endhint %}

#### Requirements and responsibilities

* **Language and ticket ownership**\
  Contacts must communicate in English. They own ticket updates and follow-ups.
* **Portal access and ticket submission**\
  Named Support Contacts can access the Customer Support Portal.\
  Only **Primary** and **Backup** contacts can open new tickets.
* **Administrative access**\
  Troubleshooting often requires admin access or elevated permissions.\
  Primary contacts should be able to approve or perform sensitive actions.
* **Keep contacts current**\
  Tell Support when a contact leaves or changes roles.\
  This ensures access is revoked and tickets are reassigned.
* **Single point of contact**\
  Named Support Contacts should not forward requests from others.\
  They act as the Support liaison for their organization.
* **Service announcements**\
  Primary contacts automatically receive critical announcements.

#### What each access level can do

| Task                    | Primary and Backup | Other Named Support Contacts | Anonymous |
| ----------------------- | ------------------ | ---------------------------- | --------- |
| Submit a ticket         | Yes                | No                           | No        |
| Knowledge Base          | Yes                | Yes                          | Limited   |
| Best practices          | Yes                | Yes                          | Yes       |
| Downloads               | Yes                | Yes                          | No        |
| Product documentation\* | Yes                | Yes                          | Yes       |
| Academy                 | Yes                | Yes                          | Yes       |

{% hint style="warning" %}
Product documentation for **Pentaho Data Quality** is restricted to registered users.

If you are licensed for this product but you do not have access, email **<support.pentaho@hitachivantara.com>**.
{% endhint %}
{% endtab %}

{% tab title="Open a ticket" %}
{% hint style="info" %}
Ticket handling is guided by three factors:

* **Subscription service level**\
  Defined in your subscription agreement.\
  If you are unsure, contact your Customer Success Manager or Sales Representative.
* **Severity level**\
  Based on business impact and user impact.
* **Environment**\
  Where the issue occurred (Production, Development, Test).

When you open a ticket, you’ll specify **severity** and **environment**. Your subscription service level is identified automatically.
{% endhint %}

#### Severity definitions

Use severity to describe **business impact** and **user impact**. Support may adjust severity based on the details provided.

* **Severity 1 (Critical)**\
  Production is down or a critical function is unavailable.\
  No workaround exists. Impact is immediate and widespread.
* **Severity 2 (High)**\
  Major functionality is impaired or performance is severely degraded.\
  A workaround exists, but it is limited or risky. Impact is significant.
* **Severity 3 (Medium)**\
  Minor functionality is affected, or you have a how-to question.\
  A workaround exists and impact is limited.

{% hint style="info" %}
For Premium and Enterprise customers, Severity 1 issues in Production may be eligible for after-hours coverage.
{% endhint %}

{% stepper %}
{% step %}

#### Do a quick self-check first

* Confirm the issue is in Pentaho, not third-party software.
* Confirm you’re on a supported version. See [Pentaho EOL Policy](https://support.pentaho.com/hc/en-us/articles/205789159-Pentaho-Product-Lifecycle-Overview).
* Try to reproduce the issue. Note if it’s consistent or intermittent.
* Search for known solutions in the Knowledge Base.

Useful resources:

* [Pentaho Product Documentation](https://docs.pentaho.com/)
* [Pentaho Academy](https://academy.pentaho.com/)
  {% endstep %}

{% step %}

#### Gather the right details

If you’re **asking a question**, include:

* What you’re trying to achieve.
* Your Pentaho product(s) and version(s).
* Environment details (OS, JVM, DB, cloud/on-prem, and so on).
* What you already tried, including exact URLs to docs or KB articles.

If you’re **reporting a product issue**, also include:

* Steps to reproduce (numbered).
* Expected result vs actual result.
* Symptoms and exact error messages.
* Logs and diagnostics. Use the [Pentaho Support Utility](/pentaho-11-installation-en/pentaho-support/pentaho-support/pentaho-support-utility) when possible.
* What changed before it started (deployments, config, certificates, upgrades).
* When it started and how often it happens.
* Business impact and urgency.
  {% endstep %}

{% step %}

#### Submit the ticket (Portal or email)

Only **Primary** and **Backup** Named Support Contacts can open new tickets.

#### Submit via portal

Use the [Pentaho Customer Portal](https://support.pentaho.com/hc/en-us) to submit a request and open a ticket.

<figure><img src="/files/nHbFI3Tc1vwP9VuWIgmd" alt=""><figcaption><p>Submit via Portal</p></figcaption></figure>

#### Submit via email

Email **<support.pentaho@hitachivantara.com>** with the details above.\
If you are a Named Support Contact, you’ll receive a confirmation email with a ticket number.

To add more information later, reply to the ticket emails.\
Include the ticket number in the email subject.

<figure><img src="/files/cp5zkGZnkcSMtbrKsaCN" alt=""><figcaption><p>Subject Line</p></figcaption></figure>
{% endstep %}

{% step %}

#### Attach visuals when it helps

Screenshots and short videos can speed up triage and reproduction. Blur or redact sensitive data before sharing.
{% endstep %}

{% step %}

#### Assigned to the support team

Support typically triages and resolves tickets in levels. If needed, Support escalates to the next level.

| Support team level                                                | Description                                                                                                                                                                         |
| ----------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Level 0: Self-help and user-retrieved information**             | Self-service resources such as FAQs, Knowledge Base articles, forums, product documentation, and troubleshooting guides. Users resolve issues independently.                        |
| **Level 1: Basic help desk resolution and service desk delivery** | Tier 1 frontline Support handles incoming inquiries, troubleshoots common problems, and provides known fixes and standard workarounds.                                              |
| **Level 2: Advanced technical support**                           | Tier 2 handles issues that Level 1 cannot resolve. This often requires deeper product knowledge and more advanced troubleshooting. Tickets may be escalated to Level 3.             |
| **Level 3: Expert product and service support**                   | Tier 3 specialists handle complex or product-specific issues. They may work with engineering on defects, enhancements, or unique requirements. Tickets may be escalated to Level 4. |
| **Level 4: Product and engineering development**                  | Engineering investigates and implements code-level fixes for the most complex issues. This level is typically involved when product changes are required.                           |
| {% endstep %}                                                     |                                                                                                                                                                                     |
| {% endstepper %}                                                  |                                                                                                                                                                                     |

#### What happens after you submit a ticket

**During business hours**

* Tickets are assigned to the next available engineer for the product area.
* The Support engineer may call you to confirm details and start quickly.
* If you cannot be reached by phone, Support contacts you in the ticket.

**Outside business hours and global holidays**

* Tickets submitted after local business hours are picked up the next business day.
* Tickets submitted on global holidays are picked up the next business day.
* For **Premium or Enterprise Severity 1 issues**, the on-call engineer responds after hours.
  {% endtab %}

{% tab title="Support hours" %}

#### Support hours

* Standard coverage: **9:00 AM to 5:00 PM, Monday to Friday** (local Support office hours).
* **24/7/365** support is available for **Premium and Enterprise** customers for **Severity 1 issues in Production**. It can also be arranged in advance for events like migrations or go-live dates.

#### Response times

Initial response targets depend on:

* Your **subscription service level**
* The ticket **severity**
* The affected **environment** (Production vs non-Production)

| Severity Level          | Enterprise       | Premium          | Starter & Standard |
| ----------------------- | ---------------- | ---------------- | ------------------ |
| **24/7/365 Production** | Yes              | Yes              | No                 |
| Severity 1              | 1 business hour  | 1 business hour  | 4 business hours   |
| Severity 2              | 2 business hours | 2 business hours | 1 business day     |
| Severity 3              | 4 business hours | 4 business hours | 2 business days    |

{% hint style="info" %}
Response-time targets vary by contract.\
If you need your specific SLA, contact your Customer Success Manager or ask Support in your ticket.

Enterprise Support includes up to **4 hours per week** of Customer Success time (Solution Architects, Customer Success Managers (CSM), or Support Account Managers (SAM)).

This time can be used to:

* Run best-practice sessions with your Named Support Contacts.
* Discuss integration techniques, solution design, and architecture.
* Discuss implementation strategies and upgrade approaches.
* Review best practices and performance tuning.
* Coordinate sessions with Pentaho subject matter experts.
* Troubleshoot issues in your systems or create solution replicas, when technically possible.

Weekly time does not accrue or roll over. You can increase the allocation by written agreement.
{% endhint %}
{% endtab %}

{% tab title="Escalate a ticket" %}
{% hint style="info" %}

#### Escalate a ticket

If the ticket priority has changed or expectations are not met, a **Primary contact** can escalate by emailing **<escalation.pentaho@hitachivantara.com>**. Include the ticket number in the subject line.

The ticket owner continues working with you on impact and next steps. Support also opens an escalation ticket with a Support management team member.
{% endhint %}
{% endtab %}

{% tab title="Defects & Enhancement Requests" %}
{% hint style="info" %}

#### Defects and enhancement requests

If Support reproduces the issue and suspects a defect, the support engineer links your ticket to a Jira issue in Product Engineering. Support changes the ticket status to **Defect** and shares major updates. Your ticket stays open while the defect is investigated.

If the request is an enhancement, Support may create a Jira enhancement. Enhancements do not have a committed implementation timeline. After requirements are captured, Support closes the ticket. Open a new ticket later if you need a status update.
{% endhint %}

<figure><img src="/files/d0m1HbuaeCmCjNHd31Uo" alt=""><figcaption><p>Ticket Status</p></figcaption></figure>

#### Automated closure timeline

| Timeframe                         | What happens                                                                                                                 |
| --------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| 24 hours after status **Pending** | You get an email that Support needs your response.                                                                           |
| 7 days after status **Pending**   | You get an email that Support has waited 7 days. Support may consider the issue resolved. Reply to keep the ticket **Open**. |
| 14 days after status **Pending**  | The system sets the ticket to **Solved** and says it will close in 7 days. A feedback survey is emailed within 24 hours.     |
| 21 days after status **Pending**  | The system sets the ticket to **Closed**. Closed tickets cannot be reopened. Submit a new ticket.                            |

{% hint style="info" %}
The automated closure process is suspended for escalated tickets and for tickets linked to a defect. You also receive a final email at ticket closure inviting you to complete a customer satisfaction survey.
{% endhint %}

{% hint style="warning" %}

#### Exclusions from Support

Pentaho does not provide support for issues caused by:

* Hardware, equipment, or programs not covered by your contract.
* Product versions not obtained through Pentaho Support.
* Use of non-general availability releases in Production (not marked as general availability (GA)).
* Causes outside Pentaho’s control (for example, floods, fires, or power loss).
* Failures outside Pentaho software (for example, databases, web servers, or hardware).
* Not following operating instructions in product documentation.
* Modifications, enhancements, or customizations performed by anyone outside Pentaho.
* Installation, configuration, management, and operation of your applications.
* APIs, interfaces, web services, or data formats not included with the product.
* Third-party products, except those provided by Pentaho, and those used only to support intended interfaces or functionality.
  {% endhint %}
  {% endtab %}
  {% endtabs %}


# Pentaho Support Utility

{% hint style="info" %}
The Pentaho Support Utility collects environment and configuration details into a single zip file that you can attach to your support ticket. This gives the support team immediate access to the background information they need to diagnose and resolve your issue, reducing back-and-forth requests for additional files and minimizing the time you spend gathering data.

The utility is supported on Pentaho versions 7.x, 8.x, 9.x, and 10.x. For a complete list of supported software and hardware, refer to the Components Reference in Pentaho Documentation.

Three tools are available for running the utility:

**Pentaho Server Plugin**: provides a UI within the Pentaho Server to gather information from the currently running server instance.

**Pentaho Data Integration Plugin**: operates within Spoon to collect details from your PDI environment.

**Command Line Utility**: offers a standalone option for gathering information from either the Pentaho Server or PDI without requiring a running application.

The server and PDI plugins are the preferred methods for running the utility. Use the command line tool only when Pentaho fails to start or when organizational policies prohibit plugin installation.

If any individual collector fails during execution, the remaining collectors will continue to run and the resulting report will still provide value to the support team. You can identify failed collectors by reviewing the utility logs or checking for .failed files within the output zip.

Some collected files and configuration details may contain passwords. While the utility attempts to remove all passwords during collection, Pentaho cannot guarantee complete removal in every case - particularly for files such as kettle.properties.
{% endhint %}

1. Download the Pentaho Support Utility.

{% embed url="<https://hcpanywhere.hitachivantara.com/userportal/?v=4.6.1#/shared/public/8eUdzjLlHa7j5R3R/245ce549-c03a-4ee7-adae-d5ea655af8ad>" %}

<figure><img src="/files/nAf9rzfIwbfXSv0e1COl" alt=""><figcaption><p>Support Utilities</p></figcaption></figure>

2. Select Pentaho Product:

{% tabs %}
{% tab title="Pentaho Server" %}
**To install the Pentaho Server Plugin:**

1. Download: `pentaho-support-utility-server-plugin.zip` file.
2. Unzip this file to: `pentaho-server/pentaho-solutions/system`.

```bash
cd
cd ~/Downloads/'Support Utility'/
unzip pentaho-support-utility-server-plugin-1.0.1.zip -d /opt/pentaho/server/pentaho-server/pentaho-solutions/system
```

{% hint style="info" %}
If needed, configure the [password removal regular expression](https://support.pentaho.com/hc/en-us/articles/360021624372-How-to-Install-and-Use-Pentaho-Support-Utility#PRCF).
{% endhint %}

3. Restart the Pentaho Server.

```sh
cd
cd /opt/pentaho/server/pentaho-server
./stop-pentaho.sh
```

```sh
cd
cd /opt/pentaho/server/pentaho-server
./start-pentaho.sh
```

{% endtab %}
{% endtabs %}


# Pipeline Designer

{% hint style="info" %}

#### Pipeline Designer

**What is Pipeline Designer?**

Pipeline Designer is a web-based interface that lets you design, execute, and manage data integration workflows directly in your browser. It supports a wide range of database connections, advanced transformation steps, and robust execution monitoring capabilities.

**Installation and Updates**

Pipeline Designer is installed by default when you install Pentaho Server. When updates are available, you can use the Plugin Manager to update the plugin. For details, see Update plugins.

**Note on Projects**

Projects appear on the Pipeline Designer main page but are not currently supported in Pipeline Designer. You can use Projects in the Pentaho Data Integration client instead. For details, see Organizing data integration with projects.

**Design and XML Views**

Pipeline Designer offers two ways to work with transformations and jobs. You can build workflows in the Design View using a visual interface, or review and edit the raw XML code in the XML View.

**ETL Workflow Concepts**

Pipeline Designer uses a workflow approach as the foundation for transforming your data. Workflows are built using individual steps that you connect together to create transformations and jobs. Each step is joined by a "hop" which passes data flow from one step to the next.

{% endhint %}

<figure><img src="/files/RiDME9d1QjV7XV5Qg9iM" alt=""><figcaption><p>Pip[eline Designer</p></figcaption></figure>

{% hint style="info" %}
**Managing Your Work**

The Pipeline Designer main page provides tools to manage your transformations and jobs. You can mark items as favorites, download them as files, create duplicates, move them between folders or to the trash, rename them, and view their details.

**Working with Transformations**

Transformations perform the actual ETL work. You create and configure transformations to define specific data processing tasks, then run them as part of your workflow. After execution, you can analyze the results to explore the data or identify improvements and problems.

**Working with Jobs**

Jobs orchestrate your ETL activities by coordinating multiple transformations and other tasks. You create and configure jobs to control the execution sequence, then run them to automate your data integration processes. Results analysis helps identify areas for improvement or troubleshooting.

**Stopping Execution**

Pipeline Designer offers two different methods for stopping running transformations or jobs. The method you choose depends on the specific processing requirements of your ETL task.

**Available Steps**

Pipeline Designer includes extensive libraries of both transformation steps and job steps. These steps extend and expand the functionality available for building your data integration workflows.
{% endhint %}

***


# Concepts & Terminology

{% hint style="info" %}

#### Concepts & Terminology

{% endhint %}

{% tabs %}
{% tab title="Projects" %}

{% endtab %}

{% tab title="Transformations" %}
{% hint style="info" %}

#### Transformations

Transformations are data flows built from interconnected nodes that process and transform data. These files use a .ktr extension and represent a directed graph of data transformation logic.

**Nodes** are the core building blocks that perform specific tasks like reading files, filtering rows, or writing to databases. Pipeline Designer provides numerous nodes organized by function (Big Data, Input, Output, Transform, etc.). You can add nodes to the canvas by dragging them from the Design pane or double-clicking. Each node runs in its own thread when the transformation executes.
{% endhint %}

<figure><img src="/files/ZfzJxLfQ3Njtbm1h28uP" alt=""><figcaption><p>Transformation</p></figcaption></figure>
{% endtab %}

{% tab title="Hops" %}
{% hint style="info" %}

#### Hops

**Hops** are the pathways connecting nodes together, shown as arrows in the interface. They define data flow direction and allow schema metadata to pass between steps. While hops may appear to create a sequence, all nodes in transformations are actually start in parallel - as a result, you cannot reliably set a variable in one node and use it in a subsequent downstream node within the same transformation. When data reaches a node with multiple outputs, it can either be copied to all destinations or distributed among them see [Data Movement](#data-movement).
{% endhint %}

**Transformations**

<figure><img src="/files/tatoFqWZkgYopKPsAWq9" alt=""><figcaption><p>Draw a Hop between Nodes</p></figcaption></figure>

<figure><img src="/files/ZEGjAdrJ8HKtlKIw3xo2" alt=""><figcaption><p>Delete or Disable Hop</p></figcaption></figure>

**Jobs**

{% hint style="info" %}
**Job Hops** in jobs behave differently than in transformations - they control execution order and specify conditions determining which node runs next based on previous results. This creates a sequential flow of control rather than parallel data processing.
{% endhint %}

<figure><img src="/files/VuPg6b1UhsTGnA6pHW4F" alt=""><figcaption><p>Always begin with 'Start' node</p></figcaption></figure>

Job hop conditions are specified in the following table:

<table><thead><tr><th width="226">Condition</th><th>Description</th></tr></thead><tbody><tr><td><strong>Unconditional</strong></td><td>Specifies that the next job entry will be executed regardless of the result of the originating job entry</td></tr><tr><td><strong>Follow when result is true</strong></td><td>Specifies that the next job entry will be executed only when the result of the originating job entry is true; this means a successful execution such as, file found, table found, without error, and so on</td></tr><tr><td><strong>Follow when result is false</strong></td><td>Specifies that the next job entry will only be executed when the result of the originating job entry was false, meaning unsuccessful execution, file not found, table not found, error(s) occurred, and so on</td></tr></tbody></table>
{% endtab %}

{% tab title="Jobs" %}
{% hint style="info" %}

#### Jobs

Jobs are workflow models that coordinate ETL activities, resources, and dependencies. Unlike transformations that focus on data flow, jobs orchestrate entire processes. A typical job might download FTP files, verify database tables exist, execute transformations to load data, and send error notifications if issues occur.

**Job Nodes** are the building blocks of jobs. The same job node can be reused multiple times with different configurations. Jobs use a .kjb extension.
{% endhint %}

<figure><img src="/files/BL2PPg2OmiS36a81pE65" alt=""><figcaption><p>Jobs</p></figcaption></figure>

1.

{% endtab %}

{% tab title="Notes" %}
{% hint style="info" %}

## Adding Notes to Transformations and Jobs

Notes help document your transformations and jobs by explaining structure, design decisions, business rules, dependencies, and other important aspects for yourself and your team.
{% endhint %}

x

{% hint style="info" %}
**Working with Notes**

**Adding a note:** Click the Add Note icon in the canvas toolbar. Enter your note content in the dialog box, optionally customize the font, color, and shadow styling, then click Save to place it on the canvas.

**Editing a note:** Hover over any note to reveal Delete and Edit icons. Click Edit to modify the note's content or formatting, then save your changes.

**Repositioning a note:** Simply click and drag the note to any location on the canvas.

**Deleting a note:** Hover over the note and click the Delete icon that appears above it.

Notes are a simple but effective way to maintain clear documentation directly within your PDI workflows, making them easier to understand and maintain over time.
{% endhint %}
{% endtab %}

{% tab title="Data Movement" %}
{% hint style="info" %}

#### Data Movement

Hops connect nodes in transformations & jobs, with arrows indicating the direction of data flow.&#x20;

**Important Constraints**

**Loops:** Transformations do not allow loops because Spoon relies on previous steps to determine field values, which could cause endless loops. Jobs do allow loops since they execute sequentially, but you must avoid creating endless loops.

**Mixed row layouts:** Transformations cannot mix rows with different layouts (such as combining table inputs with varying field counts). Mixed layouts cause failures when fields are missing or data types change unexpectedly. The trap detector warns you at design time when steps receive mixed layouts.

**Managing Hop Behavior**

Click on the three dots that appear when you hover over a node. Select "Data Movement" to specify how data is handled across multiple outgoing hops- either copied, distributed, or load balanced. You can also enable or disable individual hops, useful for testing purposes.
{% endhint %}

1. Hover over the node > click on the 3 dots.

<figure><img src="/files/cb6ZZoFxeoaB9Gffia2T" alt=""><figcaption><p>Data Movement</p></figcaption></figure>

2. From here you can Change the Number of Copies & set Data Movement.

<figure><img src="/files/GMdskDRU0fhdgMFmH8Qn" alt=""><figcaption><p>Set Numbr of Copies &#x26; Data Movement</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Hello World

Simple Transformation ..

{% hint style="warning" %}

#### Workshop - Hello World

In this hands-on workshop, you'll learn the core mechanics of building transformations in Pipeline Designer. You'll work with nodes, hops, and notes to create a simple Transformation & Job.

**What You'll Accomplish:**

* Create a new transformation from scratch
* Add and configure transformation nodes (Generate Rows and Dummy)
* Add and configure job nodes (Start, Transformation, Success)
* Connect nodes using hops to define data flow
* Add notes to document your transformation
* Preview data to verify your logic
* Interpret the execution metrics
* Understand the Logging and Step Metrics tabs for monitoring

By the end of this workshop, you'll have completed your first working transformation & job and established a repeatable development workflow: add nodes, configure properties, preview data, connect with hops, and run.&#x20;

**Prerequisites:** Pipeline Designer installed and configured

**Estimated Time:** 10 minutes
{% endhint %}

1. Log into Pentaho 11 User Console (PUC).

<figure><img src="/files/4JDcZdItKY5LOKYfFqYJ" alt=""><figcaption></figcaption></figure>

2. Select: Pipeline Designer.

<figure><img src="/files/5PPEhgP7DdEpFAtsgxuj" alt=""><figcaption><p>Pipeline Designer</p></figcaption></figure>

{% tabs %}
{% tab title="Tour of UI" %}
{% hint style="info" %}

#### Tour of Pipeline Designer UI

For folks familiar with Data Integration the key concepts & features should be familiar ..&#x20;
{% endhint %}

1. Select: Create Transformation in Pipeline Designer.

<figure><img src="/files/v1sFfzgcWE3jJRRueEs1" alt=""><figcaption><p>Pipeline Designer UI</p></figcaption></figure>

<table><thead><tr><th width="98">Feature</th><th>Description</th></tr></thead><tbody><tr><td><div><figure><img src="/files/0kyjBmx3Jq51SOsG8F8L" alt=""><figcaption></figcaption></figure></div></td><td>Collapse the left menu.</td></tr><tr><td><div><figure><img src="/files/cgvUaSZQ36t4UkWxtBFw" alt=""><figcaption></figcaption></figure></div></td><td>Switch bewteen Design &#x26; View mode.</td></tr><tr><td><div><figure><img src="/files/2A5W3jFko3Qv1qZlDRO3" alt=""><figcaption></figcaption></figure></div></td><td>Design Panel. You can also 'Search' for Nodes.</td></tr><tr><td><div><figure><img src="/files/WDMZnpRUeWtnNMQj6iFx" alt=""><figcaption></figcaption></figure></div></td><td>Add a Note</td></tr><tr><td><div><figure><img src="/files/YP9oKUyRiKygVzOWDiE6" alt=""><figcaption></figcaption></figure></div></td><td>Reset the current Transformation. It will clear all Nodes &#x26; connections. Cannot be undone.</td></tr><tr><td><div><figure><img src="/files/uy7SPgKYzM79WZibE6mn" alt=""><figcaption></figcaption></figure></div></td><td>RUN the Transformation. The drop-down Run Options enable: Run Configuration, Logging Level, Parameters &#x26; Variables</td></tr><tr><td><div><figure><img src="/files/17FkIPLdDuSmllXl4d03" alt=""><figcaption></figcaption></figure></div></td><td>Pause RUN.</td></tr><tr><td><div><figure><img src="/files/AqHxT4etKknZaGASAAO8" alt=""><figcaption></figcaption></figure></div></td><td>Stop RUN. </td></tr><tr><td><div><figure><img src="/files/T5d4ueEIFTEAj431LuUB" alt=""><figcaption></figcaption></figure></div></td><td>View the Logging, Preview Data &#x26; Step Metrics panel.</td></tr><tr><td><div><figure><img src="/files/uOn7aLBRFKxeDgQo63I5" alt=""><figcaption></figcaption></figure></div></td><td>View the Kettle Status.</td></tr><tr><td><div><figure><img src="/files/IqQYdoJqMj4QUMuryrTp" alt=""><figcaption></figcaption></figure></div></td><td>View the Transformation Properties.</td></tr><tr><td><div><figure><img src="/files/NzWUsuqUdhsAz4WATuAo" alt=""><figcaption></figcaption></figure></div></td><td>Zoom in or Out of the Canvas.</td></tr><tr><td><div><figure><img src="/files/y4bBovyjL8OFbINzUmdv" alt=""><figcaption></figcaption></figure></div></td><td>Lock the Canvas.</td></tr><tr><td><div><figure><img src="/files/Eor2oF1MLDSNvzpqzSuZ" alt=""><figcaption></figcaption></figure></div></td><td>Switch between Design &#x26; XML view.</td></tr></tbody></table>
{% endtab %}

{% tab title="Transformation" %}
{% hint style="info" %}

#### Transformation - Hello World

This workshop assumes you have some experience of using Data Integration. The intended outcome is to famialize yourself with the new features and options in the UI.&#x20;
{% endhint %}

1. Collapse the left-hand menu.
2. In pipeline Designer, drag & drop the Generate Rows node from Input onto the Canvas.

<figure><img src="/files/yYsRHbhJaHPGxNLjVzXh" alt=""><figcaption><p>Drag &#x26; Drop - Generate Rows</p></figcaption></figure>

3. Drag & Drop the Dummy node from Flow onto the Canvas.

<figure><img src="/files/u5wNJCSXOwWVwR3qsz0L" alt=""><figcaption><p>Drag &#x26; Drop - Dummy</p></figcaption></figure>

4. Create a Hop - place your cursor over the Generate rows output. It will turn to a +. Hold down the left hand mouse button and drag the Hop to the Dummy input.

<figure><img src="/files/TdWbU2kEkVOjgi2i7ldi" alt=""><figcaption><p>Create a Hop</p></figcaption></figure>

5. Double-clcik on the Generate rows node to configure

<figure><img src="/files/W55M80ODyDxJptLdedtJ" alt=""><figcaption><p>Configure Generate rows</p></figcaption></figure>

6. Enter the following details:

<table><thead><tr><th width="159">Setting</th><th>Value</th></tr></thead><tbody><tr><td>Limit</td><td>Keep default value of 10 records / rows.</td></tr><tr><td>Name</td><td>message</td></tr><tr><td>Type</td><td>String</td></tr><tr><td>Value</td><td>Hello World</td></tr></tbody></table>

<figure><img src="/files/uCuZ8608hzlAAbUKQ00n" alt=""><figcaption><p>Configure - Generate rows</p></figcaption></figure>

7. Click Preview. Enter the number of Preview rows.

<figure><img src="/files/cVMQ2Z9CBePkrXazoubJ" alt=""><figcaption><p>Preview data</p></figcaption></figure>

8. Click Cancel > Save.
9. Add a Note - you can also edit the Style.

<figure><img src="/files/PqX8stFcfcBdxmA6OFvu" alt=""><figcaption></figcaption></figure>

10. Click: Save.

<figure><img src="/files/eEwxcBN2a6gACxIHE9ht" alt=""><figcaption><p>Add Note</p></figcaption></figure>

11. Before we RUN the transformation, Save as.

<figure><img src="/files/x5IjkaQ4OuC8lEc7gO4A" alt=""><figcaption></figcaption></figure>

12. Create a Folder: Data Pipelines in 'public'.
13. Save the transformation.

<figure><img src="/files/PuNkWsb6LfiEAcflCuJm" alt=""><figcaption><p>save transformation - gr:hello-world</p></figcaption></figure>

***

{% hint style="info" %}

#### Run the Transformation

{% endhint %}

1. Click: Run.

<figure><img src="/files/lejgqvvgmgzFaRH3OyNE" alt=""><figcaption><p>Run the transformation</p></figcaption></figure>
{% endtab %}

{% tab title="Job" %}
{% hint style="info" %}

#### Job - Hello World

{% endhint %}

1. From the drop-down box, 'Create new' > Job.
2. Drag & drop the Start, Transformation & Success nodes onto the  Canvas.
3. Create the required hops. Place your cursor over the hop for more information.

<figure><img src="/files/P2pxkUsc6bws7G9F4Nfw" alt=""><figcaption><p>Hop information</p></figcaption></figure>

4. Click: Save as and Save the Job in the same folder as the previous hello-world.ktr

<figure><img src="/files/k93wQ73jj6dFiBZvfRvq" alt=""><figcaption><p>Save </p></figcaption></figure>

5. Double click on the Transformation node to configure.

<figure><img src="/files/r5j5QW1IPGK0sWPFwDm5" alt=""><figcaption></figcaption></figure>

6. Click: Select variable to insert and select: ${Internal.Entry.Current.Directory}/&#x20;
7. Add the name of the transformation - hello-world
8. Click: Save.
9. Run

x

x

x

x
{% endtab %}
{% endtabs %}


# Pentaho Data Integration

Become an expert in creating automated data pipelines ..

{% hint style="info" %}

#### **Introduction**

Pentaho Data Integration, often referred to as PDI, is a powerful tool for building your data pipelines. Its main unique features include an intuitive graphical interface, support for various data sources and destinations, extensive transformation capabilities, and robust job scheduling and automation.
{% endhint %}

{% embed url="<https://www.loom.com/share/f8661ea8298b45b281c55f12f26d7ac3?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Course Overview
{% endembed %}

{% hint style="info" %}
By the end of the workshops, you will have a comprehensive understanding of:
{% endhint %}

<details>

<summary>Pentaho Data Integration Components</summary>

</details>

<details>

<summary>Key Concepts &#x26; Terminology</summary>

</details>

<details>

<summary>Transforming &#x26; Enriching Datasets</summary>

</details>

<details>

<summary>Scaling out an Enterprise Solution</summary>

</details>

***


# Getting Started

Get up and running ..

To begin your journey with us, you will need to download and install Pentaho Enterprise - 30 day trial:

{% stepper %}
{% step %}
**Download**

Download the on-prem 30‑day Pentaho Enterprise trial:

<a href="https://pentaho.com/download/" class="button primary" data-icon="arrow-down-from-bracket">Download 30-day Pentaho Enterprise</a>
{% endstep %}

{% step %}
**Install**

Follow the Quick Start guide for a streamlined setup on Windows 11:

<a href="https://academy.pentaho.com/pentaho-11-installation-en/installation/evaluation-installation" class="button primary" data-icon="book-open">Pentaho Enterprise - Quick Start</a>
{% endstep %}

{% step %}
**Verify your install**

Launch Data Integration (GUI):

{% tabs %}
{% tab title="Windows" %}
Launch Data Integration:

```powershell
cd \
cd Pentaho/design-tools/data-integration
./spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}
Launch Data Integration:

```sh
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}
**Workshop--Data Integration**

Download the Data Integration Workshops:

```
git clone https://github.com/jporeilly/Workshop--Data-Integration
```

{% endstep %}
{% endstepper %}


# Components

Components, User Interface, Configuration options ..

{% hint style="info" %}

#### Components

Understanding the architecture and components of Pentaho Data Integration (PDI) is fundamental to becoming an effective Pentaho developer and administrator. This section will familiarize you with the building blocks that make up the Pentaho Data Integration ecosystem and how they work together to deliver enterprise-grade data integration capabilities.

**What You'll Learn**

Pentaho Data Integration operates on a **client-server architecture** that separates design-time activities from runtime execution and administration. In this section, you'll explore:

* **Enterprise Components**: The server-side infrastructure that handles execution, security, content management, and scheduling
* **Client Tools**: The desktop applications used to design, test, and deploy your data integration solutions
* **Configuration Framework**: The KETTLE configuration files that control system behaviour and store critical settings
* **Repository Management**: How PDI manages versioning, collaboration, and content organization
* **Database Connectivity**: The process of integrating JDBC drivers to connect to various data sources
  {% endhint %}

<figure><img src="/files/jGXpUwq984srpfmtAXJG" alt=""><figcaption><p>Pentaho Enterprise</p></figcaption></figure>

***

Browse to learn about the components:

{% tabs %}
{% tab title="1. Components" %}

<div data-full-width="true"><figure><img src="/files/TJzddyc54sEyeLo7Uk5P" alt=""><figcaption><p>Pentaho Client / Server Architecture</p></figcaption></figure></div>

{% tabs %}
{% tab title="1. Data Integration" %}
{% hint style="info" %}

#### **Data Integration**

**Spoon**

Graphical modelling environment for developing, testing, debugging and monitoring jobs and transformations.

**Designer**

Drag & Drop 'objects' to design your pipelines and workflows.

**Scheduler**

Connects to Quartz scheduler on server. Jobs and transformations must be uploaded to Repository.

**Engine**

Kettle and Spark engines available to execute jobs and transformations.

**Repository Browser**

Connects to Apache Jackrabbit content Repository, pointing to a supported database:

* PostgreSQL
* MSSQL Server
* Oracle
* MySQL
* MariaDB

**DB Explorer**

Database Explorer that enables you to conduct minimal database operations.
{% endhint %}

{% embed url="<https://docs.pentaho.com/pdia-data-integration>" %}
{% endtab %}

{% tab title="2. Pentaho Server" %}
{% hint style="info" %}

#### **Pentaho Server**

The Pentaho Server hosts Pentaho-created and user-created content. It is a core component for executing data integration transformations and jobs using the Pentaho Data Integration (PDI) Engine. It allows you to manage users and roles (default security) or integrate security to your existing security provider such as LDAP or Active Directory.
{% endhint %}

The primary functions of the Pentaho Server are:

<table data-header-hidden><thead><tr><th width="225"></th><th></th></tr></thead><tbody><tr><td><strong>Execution</strong></td><td>Executes ETL jobs and transformations using the Pentaho Data Integration engine</td></tr><tr><td><strong>Security</strong></td><td>Allows you to manage users and roles (default security) or integrate security to your existing security provider such as LDAP or Active Directory</td></tr><tr><td><strong>Content Management</strong></td><td>Provides a centralized repository that allows you to manage your ETL jobs and transformations. This includes full revision history on content and features such as sharing and locking for collaborative development environments.</td></tr><tr><td><strong>Scheduling &#x26; Monitoring</strong></td><td>Provides the services allowing you to schedule and monitor activities on the Data Integration Server from within the Spoon design environment (Quartz).</td></tr></tbody></table>

{% embed url="<https://docs.pentaho.com/install/components-reference>" %}
{% endtab %}

{% tab title="3. Carte" %}
{% hint style="info" %}

#### **Carte Server**

The Pentaho DI Carte Server is a vital component within the Pentaho data integration suite, designed to facilitate robust data processing operations. It serves as a stand-alone web server and execution environment that allows for the remote execution of ETL (Extract, Transform, Load) tasks, making it a cornerstone for managing data workflows efficiently.
{% endhint %}

<figure><img src="/files/OgdVoO2P96PZBNiWTwUn" alt=""><figcaption><p>Carte Server</p></figcaption></figure>

{% hint style="info" %}
**Simplicity and Efficiency**

Carte stands out for its straightforward and user-friendly setup, paired with a highly efficient operation that conserves resources. This makes it the ideal choice for organizations seeking to enhance their data integration workflows efficiently and with minimal operational burden.

**Remote Execution Flexibility**

With Carte, executing ETL tasks remotely becomes effortless, allowing for versatile data integration management from any location. Serving as a powerful remote ETL server, Carte can handle jobs and transformations from afar, broadening the capabilities of data integration strategies.

**Seamless Integration Capabilities**

Featuring an extensive array of built-in connectors, Carte excels in smoothly integrating with a multitude of sources and destinations, including databases and data warehouses. This capability facilitates straightforward data extraction, transformation, and loading processes across various platforms.

**Built for Scalability**

Carte is designed to grow with your needs, enabling deployment in multiple configurations such as Kubernetes, Docker, and cloud-based solutions. Its lightweight design ensures consistent performance, even as data demands increase.

**Intuitive Web Interface**

The Carte web interface offers a clean and efficient way to oversee jobs and transformations. Users gain access to real-time task updates, status reports, and comprehensive execution logs, all through a user-friendly dashboard.
{% endhint %}

{% embed url="<https://docs.pentaho.com/pdia-data-integration/advanced-topics-pentaho-data-integration-overview/use-carte-clusters>" %}
{% endtab %}

{% tab title="4. REST APIs" %}
{% hint style="info" %}

#### **PDI REST APIs**

You can use PDI's command line tools to execute PDI content from outside of Spoon. Typically, you would use these tools in the context of creating a script or a Cron job to run the job or transformation based on some condition outside of the realm of Pentaho software.
{% endhint %}

{% hint style="info" %}
**Pan**

A standalone command line process that can be used to execute transformations and jobs you created in Spoon. The data transformation engine Pan reads data from and writes data to various data sources. Pan also allows you to manipulate data.

```
./pan.sh /file:/home/[pentaho_user]/[path]/[transformation].ktr  /level:[Log Level]
```

{% endhint %}

{% hint style="info" %}
**Kitchen**

A standalone command line process that can be used to execute jobs. The program that executes the jobs designed in the Spoon graphical interface, either in XML or in a database repository. Jobs are usually scheduled to run in batch mode at regular intervals.

```
./kitchen.sh /file:/home/[pentaho_user]/[path]/[job].kjb  /level:[Log level]
```

{% endhint %}

{% embed url="<https://docs.pentaho.com/pentaho-rest-api/carte-apis-carte-server>" %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. PDI UI" %}
{% hint style="info" %}
**User Interface**

Within the UI, you can author, edit, run, and debug transformations and jobs. You can also enter license keys, add data connections, and define security (default options - Pentaho or LDAP).

The Welcome page contains useful links to documentation, community links for getting involved in the Pentaho Data Integration project, and links to blogs from some of the top contributors to the Pentaho Data Integration project.
{% endhint %}

<figure><img src="/files/xSc8U7jWLRfsNfq6lkS5" alt=""><figcaption><p>Welcome page</p></figcaption></figure>

{% hint style="info" %}
There are a few different ways to start PDI. The method that you should use depends on the way you installed Pentaho Data Integration (PDI).
{% endhint %}

<table><thead><tr><th width="255">OS: Windows / Unix</th><th>Action</th></tr></thead><tbody><tr><td>spoon.bat / spoon.sh</td><td>Starts Spoon</td></tr><tr><td>kichen.bat / kitchen.sh</td><td>Command Line for Jobs</td></tr><tr><td>pan.bat / pan.sh</td><td>Command Line for Transformations</td></tr></tbody></table>

**Launch Data Integration**

1. Run the following command `(Linux):`

```bash
cd
cd ~/Scripts
sh pentaho--platform.sh
```

{% content-ref url="/pages/E6q35xkzlO1UnaxsdPak" %}
[Configuring PDI UI](/pentaho-data-integration/data-integration/components/configuring-pdi-ui)
{% endcontent-ref %}
{% endtab %}

{% tab title="3. Configuration Files" %}
{% hint style="info" %}
**Configuration Files**

The default Pentaho Data Integration (PDI) HOME directory is the user's home directory. Here is located in the .kettle folder, are the main PDI configuration files.

* Windows C:{user}.kettle
* Linux based operating systems ($HOME/.kettle)

The directory may change depending on the user who is logged on. Thus, the configuration files that control the behaviour of PDI jobs and transformations are different from user to user.

This also applies when running PDI from the Pentaho BI Platform. When you set the KETTLE\_HOME variable, the PDI jobs and transformations can be run without being affected by the user who is logged on. KETTLE\_HOME is used to change the location of the files normally in \[user home].kettle
{% endhint %}

<table><thead><tr><th width="257">File</th><th>Description</th></tr></thead><tbody><tr><td>kettle.properties</td><td>main configuration file with global variables</td></tr><tr><td>shared.xml</td><td>list of shared artefacts</td></tr><tr><td>db.cache</td><td>database cache for metadata</td></tr><tr><td>repositories.xml</td><td>list of repositories</td></tr><tr><td>.spoonrc</td><td>settings for the UI</td></tr><tr><td>.languageChoice</td><td>language settings</td></tr></tbody></table>

{% tabs %}
{% tab title="3.1 kettle.properties" %}
{% hint style="info" %}
**kettle.properties**

The kettle.properties file is where you will find all the global variables for KETTLE. You can also set global variables that can be used in Transformations and Jobs. For example, you can define database connections, paths to files, or variables that can be used as parameters in your solution.
{% endhint %}

The kettle.properties can be edited using a Text Editor or via the Toolbar, select:

<div align="center"><figure><img src="/files/Y7Z1gUrMvSea3Ypf3kSV" alt=""><figcaption><p>kettle.properties</p></figcaption></figure></div>

{% content-ref url="/pages/pWX8VdbEPtf62rWINi5n" %}
[KETTLE Variables](/pentaho-data-integration/data-integration/components/kettle-variables)
{% endcontent-ref %}
{% endtab %}

{% tab title="3.2 shared.xml" %}
{% hint style="info" %}
**shared.xml**

A variety of objects can now be placed in a shared objects file on the local machine. The default location for the shared objects file is:

$HOME/.kettle/shared.xml

Objects that can be shared using this method include:

* Database connections
* Steps
* Slave servers
* Partition schemas
* Cluster schemas

The location of the shared objects file is configurable on the "Miscellaneous" tab of the Transformation > Settings dialog.
{% endhint %}

1. To share one of these objects, simply right-click on the object in the tree control on the left and choose share.

<figure><img src="/files/G5fVOhuUQGEFieInUEeZ" alt=""><figcaption><p>Shared Object - Connection</p></figcaption></figure>

{% hint style="info" %}
**Bold Type** indicates the Object is shared.
{% endhint %}
{% endtab %}

{% tab title="3.3 repositories.xml" %}
{% hint style="info" %}
**repositories.xml**

A variety of objects can now be placed in a shared objects file on the local machine. The default location for the shared objects file is:

$HOME/.kettle/repositories.xml
{% endhint %}

```xml
<repositories>
<repository>
<id>PentahoEnterpriseRepository</id>
<name>Pentaho</name>
<description/>
<is_default>false</is_default>
<repository_location_url>http://localhost:8080/pentaho</repository_location_url>
<version_comment_mandatory>N</version_comment_mandatory>
</repository>
</repositories>
```

{% endtab %}

{% tab title="3.4 .spoonrc" %}
{% hint style="info" %}
**.spoonrc**

Used to store preferences and program state of Spoon. Other Kettle programs do not use this file.

* General settings and defaults
* User interface settings
* The last opened transformation/job

The default location for the shared objects file is:

$HOME/.kettle/.spoonrc
{% endhint %}

```
#Kettle Properties file
#Sat Dec 16 22:49:28 GMT 2023
AskAboutReplacingDatabases=N
AutoCollapseCoreObjectsTree=Y
AutoSave=N
AutoSplit=N
BackgroundColorB=255
BackgroundColorG=255
BackgroundColorR=255
CustomParameterMergeJoinSortWarning=Y
CustomParameterMergeRowsSortWarning=Y
CustomParameterSetVariableUsageWarning=Y
...
```

{% hint style="info" %}
These options are set from the main menu: Tools -> Options
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="4. JDBC" %}
{% hint style="info" %}
**Adding JDBC Drivers**

The PDI & Pentaho Server needs the appropriate driver to connect to the database that stores your data. Your database administrator, Chief Intelligence Officer, or IT manager should be able to provide the appropriate driver. If not, you can download drivers from your database vendor's website.

The [Components Reference](https://help.hitachivantara.com/Documentation/Pentaho/9.0/Setup/Components_Reference) contains a list of drivers.

Once you have the correct driver, copy it to the following directories:

* Pentaho Server: /pentaho/server/pentaho-server/tomcat/lib/
* PDI client: data-integration/lib
  {% endhint %}

{% hint style="danger" %}
You must restart the PDI client for the driver to take effect.

There should be only one driver for your database in the directory. Ensure that there are no other versions of the same vendor's driver in this directory. If there are, back up the old driver files and remove them to avoid version conflicts.
{% endhint %}
{% endtab %}

{% tab title="5. Repository" %}
{% hint style="info" %}
**Pentaho Repository**

The Pentaho+ platform implements its repository using [**Apache Jackrabbit**](http://jackrabbit.apache.org), a fully conforming implementation of the content repository for Java technology API (JCR, specified in JSR 170 and JSR 283)

Apache Jackrabbit needs two pieces of information to set up a runtime content repository instance:

* Repository home directory\
  The filesystem path of the directory containing the content repository accessed by the runtime instance of Jackrabbit. This directory usually contains all the repository content, search indexes, internal configuration, and other persistent information managed within the content repository. Note that this is not absolutely required and some persistence managers and other Jackrabbit components may well be configured to access files and even other resources (like remote databases) outside the repository home directory. A designated repository home directory is however always needed even if some components choose to not use it. Jackrabbit will automatically fill in the repository home directory with all the required files and subdirectories when the repository is first instantiated.
* Repository configuration file\
  The filesystem path of the repository configuration XML file. This file specifies the class names and properties of the various Jackrabbit components used to manage and access the content repository. Jackrabbit parses this configuration file and instantiates the specified components when the runtime content repository instance is created.
  {% endhint %}

{% hint style="info" %}
[**Hibernate**](https://hibernate.org/) is a Java framework which is used to store the Java objects in the relational database system. It is an open-source, lightweight, ORM (Object Relational Mapping) tool.
{% endhint %}

{% hint style="info" %}
[**Quartz**](http://www.quartz-scheduler.org/) is an open source job-scheduling framework written entirely in Java and designed for use in both *J2SE* and *J2EE* applications.
{% endhint %}

***

{% hint style="info" %}
**Versioning & Comments (Dev only)**

Pentaho Data Integration (PDI) can track versions and comments for jobs, transformations, and connection information when you save them. You can turn version control and comment tracking on or off by modifying their related statements in the repository.spring.properties text file.

By default, version control and comment tracking are disabled (set to false). Best Practice: manage your ETL workflows with a 3rd party content management tool, e.g. Github; only uploading the production version into the Repository.
{% endhint %}

1. Exit from the PDI client (also called Spoon).
2. Stop the Pentaho Server.
3. Edit repository.spring.properties file.

```bash
cd
cd ~/Pentaho/server/pentaho-server/pentaho-solution/systems
nano repository.spring.properties
```

4. Edit the versioningEnabled and versionCommentsEnabled statements:

```
versioningEnabled=true versionCommentsEnabled=true
```

{% endtab %}
{% endtabs %}


# Configuring PDI UI

PDI UI configuration settings ..

{% hint style="warning" %}

#### **Workshop - Configuring PDI**

Configure Spoon before you build transformations and jobs. Set a few defaults that speed up daily work.

You will start Spoon, review the Welcome page, and update UI options.

**What You'll Accomplish:**

* Start Spoon and confirm your install.
* Find docs and community links from the Welcome page.
* Update General and Look & Feel options.
* Switch perspectives and understand what changes.

You will finish with a clean, predictable workspace. You will know where to change these settings later.

**Prerequisites:** Pentaho Data Integration installed and ready to launch

**Estimated Time:** 10 minutes
{% endhint %}

***

{% hint style="info" %}

#### **Configuring PDI UI**

Use **Tools > Options** to configure Spoon. Start with **General** and **Look & Feel**.
{% endhint %}

{% tabs %}
{% tab title="1. Welcome Page" %}
{% hint style="info" %}

#### **Welcome page**

The Welcome page has links to:

* Documentation
* Community Forum
  {% endhint %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. The PDI UI - Spoon will be displayed.

<figure><img src="/files/xSc8U7jWLRfsNfq6lkS5" alt="Spoon Welcome page"><figcaption><p>Welcome screen</p></figcaption></figure>
{% endtab %}

{% tab title="2. Modify Look & Feel" %}
{% hint style="info" %}

#### **Kettle Options**

Set a few defaults once. It reduces prompts and visual noise.
{% endhint %}

1. In Spoon, click **Tools > Options**.
2. Open the **General** tab.

<div align="center"><figure><img src="/files/x2Nz5KMDgXOKd0VO5RBl" alt="Kettle Options dialog showing General settings" width="375"><figcaption><p>kettle options - general</p></figcaption></figure></div>

* [ ] Uncheck the ‘Show tips at startup?’ checkbox.
* [ ] Uncheck the ‘Use database cache’ checkbox.
* [ ] Uncheck the ‘Show repository dialog at startup?’ checkbox.
* [ ] Uncheck the ‘Ask user when exiting?’ checkbox.

3. Open the **Look & Feel** tab.
4. Review these settings and update as needed:
   * **Grid size** and **snap** behavior.
   * Canvas look (colors, anti-aliasing).
   * **Preferred language** and **alternative language**.

<figure><img src="/files/4lGwXZ91vllbDBx0wpZK" alt="Kettle Options dialog showing Look &#x26; Feel settings"><figcaption><p>Kettle options - look &#x26; feel</p></figcaption></figure>

{% embed url="<https://docs.pentaho.com/pdia-data-integration/get-started-with-the-pdi-client-1/customize-the-pdi-client>" %}
{% endtab %}

{% tab title="3. Perspectives" %}
{% hint style="info" %}

#### **Perspectives**

Use perspectives to switch your workspace. You can move between design and scheduling views.

Switch between:

* Designing ETL jobs and transformations
* Scheduling jobs and transformations

Click the **Perspective** icon in the toolbar to switch.
{% endhint %}

<div align="center"><figure><img src="/files/147aQOgWL53J4jGfE6o2" alt="Perspective switcher in the Spoon toolbar" width="375"><figcaption><p>Perspectives</p></figcaption></figure></div>

{% hint style="info" %}
You must connect to a repository to schedule jobs and transformations.
{% endhint %}

{% embed url="<https://docs.pentaho.com/pdia-data-integration/get-started-with-the-pdi-client-1/use-the-pdi-client-perspectives>" %}
{% endtab %}
{% endtabs %}


# KETTLE Variables

The kettle.properties file contains global variables for KETTLE.

{% hint style="warning" %}

#### **Workshop - Kettle Variables**

Use variables to avoid hardcoded paths and values.

Use `kettle.properties` to store global variables for Spoon.

In this workshop, you will create a global variable. You will then use it in a transformation.

**What You'll Accomplish:**

* Access and edit the kettle.properties configuration file
* Define a global variable for jobs and transformations
* Use both variable formats (`${VAR}` and `%%VAR%%`)
* Insert variables with `Ctrl+Space`

**Prerequisites:** Pentaho Data Integration installed and configured

**Estimated Time:** 10 minutes
{% endhint %}

***

{% hint style="info" %}

#### **Global Variables - kettle.properties**

Variables can be used throughout Pentaho Data Integration, including in transformation steps and job entries. You define variables by setting them with the Set Variable step in a transformation or by setting them in the kettle.properties file in the directory.

Use variables by either retrieving them with the Get Variable step or by using metadata strings like:

* `${VARIABLE}`
* `%%VARIABLE%%`

You can mix both formats. The first is Unix-style. The second is Windows-style.

Fields that support variables show the blue `${}` icon.

Press `Ctrl+Space` to insert a variable in those fields. Hover over the icon to see help.
{% endhint %}

{% embed url="<https://www.loom.com/share/89d5419da317432eb272b583e30f8904?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
KETTLE Variables
{% endembed %}

{% file src="/files/o77ZkN95sy2DlfKXYFwh" %}

***

1. Start Pentaho Data Integration.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Select Edit -> Edit the kettle.properties file
3. Highlight the first row and right mouse click, and select the following option.

<figure><img src="/files/RVMx57NJQ5JiGf1GCYsP" alt=""><figcaption><p>Global variables</p></figcaption></figure>

4. Add a `DIR_SAMPLES` variable for your OS.

{% tabs %}
{% tab title="Windows" %}
Add this line:

```properties
DIR_SAMPLES=C:/Temp
```

{% endtab %}

{% tab title="Linux/macOS" %}
Add this line:

```properties
DIR_SAMPLES=/home/pentaho/Temp
```

{% endtab %}
{% endtabs %}

5. Save.

{% hint style="info" %}
Spoon loads `kettle.properties` on startup.

If variables do not show up, restart Spoon.
{% endhint %}

{% hint style="info" %}
You can also edit `kettle.properties` manually.

Default locations:

* Windows: `C:\Users\<username>\.kettle\kettle.properties`
* Linux/macOS: `~/.kettle/kettle.properties`

The PowerShell script uses [nano](https://github.com/okibcn/nano-for-windows) which was installed using [scoop](https://scoop.sh/)
{% endhint %}

6. Open a terminal.

{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
cd \
cd Workshop--Data-Integration\Scripts
.\edit-kettle.properties.ps1
```

{% endtab %}

{% tab title="Linux/macOS" %}

```bash
cd
cd ~/.kettle
nano kettle.properties
```

{% endtab %}
{% endtabs %}

<figure><img src="/files/FdiiVoSk4yNfvs9CeDqM" alt="" width="563"><figcaption><p>kettle.properties - Linux</p></figcaption></figure>

{% hint style="info" %}
Verify in Spoon:

1. Open any step property that shows the blue `${}` icon.
2. Press `Ctrl+Space`.
3. Search for `DIR_SAMPLES`.
4. Insert it into the field.
   {% endhint %}

<div align="center"><figure><img src="/files/rIBYtJr1YPxKnoFnbHYu" alt="" width="375"><figcaption><p>Global variable in Transformation</p></figcaption></figure></div>

{% embed url="<https://docs.pentaho.com/pdia-data-integration/data-integration-perspective-in-the-pdi-client/advanced-topics-pdi-perspective/pdi-run-modifiers/variables/kettle-variables>" %}

***


# Concepts & Terminology

Understanding the key concepts & lingo ..

{% hint style="info" %}

#### Concepts & Terminology

The Data Integration perspective allows you to create two basic workflow types:

**Transformations**

Transformations are used to describe the data flows for ETL such as reading from a source, transforming data and loading it into a target location.

**Jobs**

Jobs are used to coordinate ETL activities such as defining the flow and dependencies for what order transformations should be run, or prepare for execution by checking conditions such as, "Is my source file available?" or "Does a table exist in my database?"
{% endhint %}

### Transformations & Jobs

{% tabs %}
{% tab title="1. Transformation" %}
{% hint style="info" %}

#### **Transformation**

Transformations are the workhorses of the ETL process. They are comprised of:

**Steps**

which provide you with a wide range of functionality ranging from reading text-files to implementing slowly changing dimensions.

Steps executed in parallel.

**Hops**

help you define the flow of the data in the stream. They represent a row buffer between the Step Output and the next Step Input, as illustrated in the below Transformation. Data flows from the Text file input step to Filter rows to Sort Rows, finally to Table output.
{% endhint %}

{% embed url="<https://pentaho-public.atlassian.net/wiki/spaces/EAI/overview?homepageId=363267360>" %}

<div data-full-width="true"><figure><img src="/files/jEAU7sQFwOXrlyI3aXOn" alt="" width="563"><figcaption><p>Steps &#x26; Hops = Transformation</p></figcaption></figure></div>

***

{% hint style="info" %}

#### **Steps**

There are some key characteristics of Steps:

* Step names must be unique in a single Transformation
* Virtually all Steps read and write rows of data (exception Generate rows)
* Most Steps can have multiple outgoing hops. These can be configured to either copy or distribute the data. Copy ensures all Steps receive a copy of the row of data; Distribute sends the data in a round robin fashion to each of the Steps.
* Steps run in their own thread. It’s possible to run multiple copies of the Step, for performance tuning, each in their own thread.
* All Steps are executed in parallel, so it’s not possible to define an order of execution.
  {% endhint %}

In addition to Steps, Hops, and Notes enable you to document the Transformation.

<div align="left"><figure><img src="/files/ogTCE6HOyY5wtIUulo2h" alt="" width="563"><figcaption><p>copy rows</p></figcaption></figure></div>

<div align="left"><figure><img src="/files/IbU3lgNRdnMBtbVJhU2h" alt="" width="563"><figcaption><p>distribute rows</p></figcaption></figure></div>

<figure><img src="/files/CcEZdAgJ8Epx7MBdsTh8" alt=""><figcaption><p>Steps</p></figcaption></figure>
{% endtab %}

{% tab title="2. Parallelism" %}
{% hint style="info" %}

#### **Parallelism**

When a transformation starts, all steps start at the same time. The hop is configured as a buffer, with generally a 10k row set.

The flow for the data stream occurs when the first step has initialized, started reading the first row sets, then writing them into the hop (10 k buffer). The row sets are then read by the next step, while the first step is still reading and writing row sets into the stream, and the second step outputs into the stream for the next step, and so on.. The buffer size can be set in Miscellaneous tab, in the Transformation properties panel.
{% endhint %}

<figure><img src="/files/N9VGarvSXgXoAfD79f3S" alt="" width="563"><figcaption><p>parallelism</p></figcaption></figure>

{% hint style="info" %}

#### **Adjusting the Queue Size**

When trying to optimize performance, you may want to adjust the input/output queue size. Especially if you have a lot of RAM available. The queue size is configured as the “Nr of rows in rowset” in the transformation settings and applies to all transformation steps. Increasing it might finish the opening steps of a transformation more quickly, thus freeing up CPU time for the subsequent steps.
{% endhint %}

<div align="center"><figure><img src="/files/EutWa9CsGNTZBmNmBt6R" alt="" width="563"><figcaption><p>Changing buffer row set</p></figcaption></figure></div>
{% endtab %}

{% tab title="3. Data Types" %}
{% hint style="info" %}

#### **Data Types**

PDI data types map internally to Java data types, so the Java behavior of these data types applies to the associated fields, parameters, and variables used in your transformations and jobs.
{% endhint %}

The following table describes these mappings:

<table><thead><tr><th width="176.66666666666666">PDI Data Type</th><th width="149">Java Data Type</th><th>Description</th><th>Example</th></tr></thead><tbody><tr><td>BigNumber</td><td>BigDecimal</td><td>An arbitrary unlimited precision number.</td><td>3.141592653589793238462643383279502884197169399375105820974944</td></tr><tr><td>Binary</td><td>Byte[]</td><td>An array of bytes that contain any type of binary data.</td><td>An image file or a compressed file can be stored as Binary data</td></tr><tr><td>Boolean</td><td>Boolean</td><td>A boolean value <code>true</code> or <code>false.</code></td><td>A boolean value true or false</td></tr><tr><td>Date</td><td>Date</td><td>A date-time value with millisecond precision.</td><td>2023-10-20T10:48:51.123</td></tr><tr><td><p><mark style="color:red;">Hierarchical -</mark></p><p><mark style="color:red;">EE Plugin 9.5+</mark></p></td><td>BinaryTree</td><td>Data items that are related to each other by hierarchical relationships</td><td>A family tree</td></tr><tr><td>Integer</td><td>Long</td><td>A signed long 64-bit integer.</td><td>42</td></tr><tr><td>Internet Address</td><td>InetAddress</td><td>An Internet Protocol (IP) address.</td><td>192.168.0.1</td></tr><tr><td>Number</td><td>Double</td><td>A double precision floating point value (64bits).</td><td>2.7182818284590452353602874713526624977572470936999</td></tr><tr><td>String</td><td>String</td><td>A variable unlimited length text encoded in UTF-8 (Unicode).</td><td>“Hello world!”</td></tr><tr><td>Timestamp</td><td>Timestamp</td><td>Allows the specification of fractional seconds to a precision of nanoseconds.</td><td>2023-10-20T10:48:51.123456789</td></tr></tbody></table>
{% endtab %}

{% tab title="4. Jobs" %}
{% hint style="info" %}

#### **Jobs**

In a PDI process, jobs orchestrate other jobs and transformations in a coordinated way to realize our business process:
{% endhint %}

{% hint style="info" %}

#### **Job Entries**

Represent the different tasks or processes that need to be executed as part of the job. Job entries can include Transformations, shell scripts, database operations, file operations, and more. Each job entry performs a specific task and can be configured with various options and parameters.

Entries executed sequentially.
{% endhint %}

<figure><img src="/files/zidnZsOgVTO9fXm5DCD5" alt=""><figcaption><p>Job Entries</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Hello World

Simple transformation to illustrate key concepts ..

{% hint style="warning" %}

#### Workshop - Hello World

Build a minimal transformation in Spoon. Use steps, hops, and notes. Preview data and review execution metrics.

**What you’ll do**

* Create a transformation
* Add and configure **Generate Rows** and **Dummy**
* Connect steps with hops
* Add a note to document the flow
* Preview data from a step
* Run the transformation and review results

**Prerequisites:** Pentaho Data Integration installed and configured

**Estimated time:** 10 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/cb9fa033ddcf46c6b7d17c532c16ac66?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
What is a Transformation?
{% endembed %}

***

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

<figure><img src="/files/loS5a4lF51rXARcNKI8Z" alt="" width="375"><figcaption><p>hello world.tr</p></figcaption></figure>

{% file src="/files/mkALCJT1B7bbbTJUrbOb" %}

{% tabs %}
{% tab title="1. Generate Rows" %}
{% hint style="info" %}

#### **Generate Rows**

Generate Rows outputs a specified number of rows. By default, the rows are empty. You can also generate static fields for test data. For example, generate 12 rows for 12 months.

Generate Rows is also useful as a single-row “starter” step.
{% endhint %}

1. Start Pentaho Data Integration (Spoon).

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. In the **Design** tab, expand the `Input` category.
3. Drag **Generate Rows** onto the canvas.

{% hint style="info" %}
Tip: You can also search for `Generate Rows`.
{% endhint %}

<div align="center"><figure><img src="/files/IXTouWAqQL6fG0HgsgKh" alt="" width="563"><figcaption><p>Generate rows</p></figcaption></figure></div>

4. Double-click **Generate Rows** to open the step properties.

<figure><img src="/files/CsNI3iKE2qYYQDPoSDU7" alt="" width="563"><figcaption><p>Generate rows - settings</p></figcaption></figure>

Ensure the following details are configured:

| **Step name** | gr\_hello-world |
| ------------- | --------------- |
| **Limit**     | 10              |
| **Name**      | message         |
| **Type**      | string          |
| **Value**     | hello world     |

{% hint style="info" %}
Before you close the dialog, preview the data.
{% endhint %}

5. Select **Preview**. The **Enter preview size** dialog opens.

<figure><img src="/files/5ZKGGptBW498f2uKk0Fg" alt="" width="375"><figcaption><p>Preview rows</p></figcaption></figure>

6. In **Enter preview size**, select **OK**.
7. Verify the 10 rows. Then select **OK** to close the preview dialog.
8. Select **OK** to close the **Generate Rows** dialog.
   {% endtab %}

{% tab title="2. Dummy" %}
{% hint style="info" %}

#### **Dummy**

The Dummy step does not process records. Use it as a placeholder during development. It is handy when you need a second step to connect.
{% endhint %}

1. In the **Design** tab, expand the `Flow` category.
2. Drag **Dummy** onto the canvas.

<div align="center"><figure><img src="/files/vX5f3hFE4hrlQk4m9SSY" alt="" width="375"><figcaption><p>dummy step</p></figcaption></figure></div>
{% endtab %}

{% tab title="3. Hops, Annotations, etc" %}
{% hint style="info" %}

#### **Hops**

Hops define row flow between steps. PDI buffers rows between steps as the transformation runs.
{% endhint %}

1. Select the `gr_hello-world` step.
2. Hold down the Shift key.
3. Drag and drop the hop onto the Dummy step.
4. Release the Shift key.

***

**Add a note**

1. Right-click anywhere on the Spoon canvas.
2. Select **New note**.

<div align="left"><figure><img src="/files/TzVqv5WMVmC5NmcJQq2g" alt=""><figcaption><p>Notes</p></figcaption></figure> <figure><img src="/files/bDztZRAqaULGC2DxZ9iM" alt=""><figcaption><p>Notes - Style</p></figcaption></figure></div>

***

**Transformation properties**

To view the transformation properties:

1. Double-click anywhere on the canvas.

<div align="center"><figure><img src="/files/aiJ2m5t9FJ8RAJjtr36g" alt="" width="563"><figcaption><p>transformation properties</p></figcaption></figure></div>

{% hint style="info" %}
Tip: Add details in **Extended description**.
{% endhint %}
{% endtab %}

{% tab title="4. Run" %}
{% hint style="info" %}

#### Run the transformation

Run the transformation locally.
{% endhint %}

1. In Spoon, select **Action** > **Run this transformation**.

{% hint style="info" %}
You can also select **Run** in the toolbar.

The **Execute a transformation** window opens. For this workshop, keep **Local** execution.
{% endhint %}

2. In the run dialog, open **Run options**.

<div align="center"><figure><img src="/files/SOex7kJn78020PnzezYU" alt="" width="375"><figcaption><p>Run Options</p></figcaption></figure></div>

{% hint style="info" %}
In the Run Options panel you can set:

* **Run configuration** (local, remote, or cluster)
* **Log level**
* **Automatically save** the transformation
  {% endhint %}

<div align="center"><figure><img src="/files/9qTcFeoYTWCHIlnHpydy" alt="" width="375"><figcaption><p>Automatically save transformation</p></figcaption></figure></div>

The transformation executes.

<figure><img src="/files/tflxq3QDgtwEqkVoGgMG" alt="" width="375"><figcaption><p>Green ticks indicate successful execution</p></figcaption></figure>

{% hint style="warning" %}
A green tick confirms the transformation's execution, but doesn't guarantee the success of the underlying operations.
{% endhint %}

***

{% hint style="info" %}
**Execution Results**

The Execution Results section of the window contains several different tabs that help you to see how the transformation executed, pinpoint errors, and monitor performance.
{% endhint %}

<div align="center"><figure><img src="/files/3abSQJPSwtHOWwCasPVC" alt=""><figcaption><p>Logging</p></figcaption></figure></div>

Logging tab displays logging information for each of the steps in the transformation.

<figure><img src="/files/0vghN3SGBFmq0vKIzs1h" alt=""><figcaption><p>Step Metrics</p></figcaption></figure>

{% hint style="info" %}
Step Metrics tab provides statistics for each step in your transformation including how many records were read, written, caused an error, processing speed (rows per second) and more. This tab also indicates whether an error occurred in a transformation step.
{% endhint %}

<figure><img src="/files/ve4KWGOLnLXSYvU8HgTp" alt=""><figcaption><p>Metrics</p></figcaption></figure>

{% hint style="info" %}
Metrics can help identify bottlenecks (back pressure). In this example, the transformation took 30 ms. Notice `gr_hello-world` and `Dummy` initialize at the same time. Steps run in parallel in separate threads.
{% endhint %}

<figure><img src="/files/ztoxSZh5BrtCXYscMNhW" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="info" %}
Preview tab displays the records.
{% endhint %}

***

{% hint style="info" %}
**Viewing the Transformation structure**

Select the **View** icon (upper-left). The tree switches to the structure of the transformation.
{% endhint %}

<div align="center"><figure><img src="/files/IMq4DCL6xRNrMWGpDIuK" alt=""><figcaption><p>View</p></figcaption></figure></div>
{% endtab %}
{% endtabs %}

***


# Logging

Set the transformation logging level ..

{% hint style="warning" %}

#### Workshop - Logging

Use logging to diagnose transformation problems. Create a controlled type mismatch error. Use log level and output to find the cause.

**What you’ll do**

* Change field metadata to trigger an error
* Run with **Basic** and **Row level** logging
* Use **Execution results** to find the failing step
* Locate the same error in `pdi.log`

**Prerequisites:** Complete the **Hello World** workshop

**Estimated time:** 5 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/868c1e73bdeb464e9ef5c2dd3c220d61?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Logging
{% endembed %}

***

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

<figure><img src="/files/loS5a4lF51rXARcNKI8Z" alt="" width="375"><figcaption><p>Logging</p></figcaption></figure>

{% file src="/files/mkALCJT1B7bbbTJUrbOb" %}

{% tabs %}
{% tab title="1. Modify Generate Rows" %}
{% hint style="info" %}
Start from the transformation you built in **Hello World**.
{% endhint %}

{% hint style="info" %}

#### **Generate Rows**

Generate Rows is a test-data step. It is a quick way to validate logging and error handling.
{% endhint %}

{% hint style="warning" %}
To create an error, change the field type for `message` from **String** to **Integer**.
{% endhint %}

1. Double-click the **Generate Rows** step.
2. Change the type for `message` to **Integer**.

<div align="center"><figure><img src="/files/Za4qHHAeBcKbC0di993V" alt="" width="375"><figcaption><p>Change data type</p></figcaption></figure></div>

3. Select **OK**.
   {% endtab %}

{% tab title="2. Run" %}
{% hint style="info" %}

#### Run the transformation

Run with a higher log level to see row-level detail.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. Set **Log level** to **Basic**.
3. Select **Run**.
4. Run again with **Log level** set to **Row level**.

<div align="center"><figure><img src="/files/WdjoimzUF2XWYEo4UPwu" alt="" width="563"><figcaption><p>Set row level debugging</p></figcaption></figure></div>

5. Select **Run**. The failing step is highlighted.

<figure><img src="/files/LDApRREPhxV1kUh5ZkZ7" alt="" width="375"><figcaption><p>Error in step</p></figcaption></figure>

6. In **Execution results**, open the **Log** tab.

<figure><img src="/files/Vm6XnlJlsWboa4ioPCIZ" alt=""><figcaption><p>Logging - error</p></figcaption></figure>

The error text is in the log output. Look for the first **ERROR** entry.

{% hint style="info" %}
Tip: Select the minus icon to show errors only.

The same error is written to `pdi.log`:

{% tabs %}
{% tab title="Windows" %}
`C:\\Pentaho\\design-tools\\data-integration\\logs\\pdi.log`
{% endtab %}

{% tab title="macOS / Linux" %}
`~/Pentaho/design-tools/data-integration/logs/pdi.log`
{% endtab %}
{% endtabs %}
{% endhint %}

<figure><img src="/files/EOgcXcFs3DF8LMUO1GFs" alt=""><figcaption><p>pdi.log - Linux</p></figcaption></figure>
{% endtab %}
{% endtabs %}

***

Next workshop: [Error Handling](/pentaho-data-integration/data-integration/concepts-and-terminology/error-handling)


# Error Handling

Handling errors in a transformation ..

{% hint style="warning" %}

#### Workshop - Error Handling

Bad data happens. Don’t fail the whole transformation because of a few rows. Route error rows to a separate stream for review and cleanup.

**What you’ll do**

* Read a CSV with a date field
* Trigger a controlled date parsing error
* Configure an error hop to capture failing rows
* Review the error metadata fields (description, field name, error code)
* Fix the date format and verify success

**Prerequisites:** Complete the **Hello World** and **Logging** workshops

**Estimated time:** 10 minutes
{% endhint %}

***

<figure><img src="/files/eqdOBPvxSiAQqWlgomiO" alt=""><figcaption><p>Error handling</p></figcaption></figure>

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/qbHHtBgSQAPhXGGQnDtE" %}

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

{% tabs %}
{% tab title="1. CSV file input" %}
{% hint style="info" %}

#### **CSV file input**

The CSV File Input step reads data from delimited text files into a PDI transformation. While this step is called CSV File Input, you can also use CSV File Input with many other separator types, such as pipes, tabs, and semicolons.

**Note:** The semicolon (;) is set as the default separator type for this step.
{% endhint %}

1. Double-click to edit the CSV file input step.

<figure><img src="/files/YiJRcLaPaZIbXKQO01YV" alt=""><figcaption><p>CSV file input</p></figcaption></figure>

2. Set the following metadata properties for: birthdate

| Fieldname | Type | Format     |
| --------- | ---- | ---------- |
| birthdate | date | yyyy/MM/dd |

{% hint style="warning" %}
If your CSV uses a different date pattern, keep `yyyy/MM/dd` for now. This mismatch is what triggers the error rows in the next step.
{% endhint %}
{% endtab %}

{% tab title="2. Error hop" %}
{% hint style="info" %}

#### Error hop

An error hop routes rows that fail in a step to a separate target step. This lets you keep processing valid rows. You also get extra error fields in the error stream.
{% endhint %}

1. Double-click the white diagonal cross on the red error hop.

<figure><img src="/files/4J94cekxgltyfD170Gym" alt=""><figcaption><p>Hop - Error handling</p></figcaption></figure>

2. Set the error field names (you can pick your own).

* **Nr of errors fieldname**: Number of errors for the row.
* **Error descriptions fieldname**: Human-readable error message.
* **Error field fieldname**: The field that caused the error.
* **Error codes fieldname**: A code you can filter or group by.
  {% endtab %}

{% tab title="3. Run" %}
{% hint style="info" %}

#### Run the transformation

Preview both streams. One contains valid rows. One contains error rows plus error metadata.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. Preview the **Dummy** step:

<figure><img src="/files/HgMqUQhlZ62FeL5KBMCH" alt=""><figcaption><p>Correct birthdate format</p></figcaption></figure>

3. Preview the **Dummy - Errors** step:

<figure><img src="/files/21KRmIJRrGKS7K5Ds1EH" alt=""><figcaption><p>Errors for incorrectly formatted birthdates</p></figcaption></figure>

4. Scroll to the end of the **Execution results** pane.

{% hint style="info" %}
Use `errorCodes` to route errors into targeted cleanup logic.
{% endhint %}

***

**Fix the format and verify**

1. Open **CSV file input** again.
2. Update the **Format** value for `birthdate` to match your CSV.

{% hint style="info" %}
Example: if your data looks like `2026-02-17`, use `yyyy-MM-dd`.
{% endhint %}

3. Run the transformation again.
4. Preview **Dummy - Errors**. You should see fewer rows, or none.
   {% endtab %}
   {% endtabs %}

***


# Projects

Configure a group of assets into a Project ..

{% hint style="info" %}

#### NEW - Data Integration Projects

NEW in Pentaho Data Integration (PDI) v11 projects help you organize ETL workflows by grouping related transformations, jobs, and metadata into logical containers. This structure brings clarity and consistency across all your environments.

**Organization and collaboration:** Projects reduce complexity by segmenting large data environments into manageable modules. Each project is self-contained, making it easy to share, move, or track in version control without losing configurations. This supports team collaboration and maintains a complete history of changes.

**Environment management:** Projects ensure your workflows run consistently across development, QA, and production. Layered configuration settings at the project, system, and repository levels clarify where each setting comes from, while reusable configurations eliminate repetitive manual edits. Configuration isolation prevents conflicts between different projects.

**Developer experience:** The Spoon UI provides dedicated menus and file trees specifically for working with projects, making navigation intuitive. Projects maintain backward compatibility, so your existing workflows continue to run without modification.
{% endhint %}

<figure><img src="/files/3IpIqsIDxD0amxdCUwOI" alt=""><figcaption><p>PDI Project</p></figcaption></figure>


# Project - Sales DWH

PDI Lifecycle Management ..

{% hint style="warning" %}

#### Workshop - Sales DWH

This comprehensive workshop guides you through setting up a professional-grade Sales Data Warehouse project using Pentaho Data Integration (PDI). You'll learn not just HOW to build ETL solutions, but WHY certain architectural decisions matter for long-term success.

**What You'll Accomplish:**

By the end of this workshop, you will be able to:

* Set up a structured PDI project with proper governance
* Integrate PDI with Git for version control
* Implement environment-agnostic configurations
* Build a reusable DI framework for logging, monitoring, and error handling
* Apply enterprise-grade best practices to your ETL projects
* Deploy solutions across multiple environments

**Target Audience:**

* ETL/Data Integration Developers
* Data Engineers transitioning to Pentaho
* Technical Leads planning large-scale DI projects
* DevOps engineers supporting PDI deployments

**Estimated Time:**

* **Full Workshop**: 2 days (16 hours)
* **Core Modules**: 1 day (8 hours)
  {% endhint %}

{% tabs %}
{% tab title="1. Challenge" %}
{% hint style="info" %}

#### Set the scene ..

Imagine you've been tasked with building a Sales Data Warehouse that needs to integrate data from multiple source systems like your CRM, ERP, and various flat files. This isn't a small project - you'll have a team of 5-8 developers working together, and the solution needs to work seamlessly across multiple environments: Development, Test, UAT, and Production.

The warehouse will require frequent updates and deployments as business requirements evolve, and it must include comprehensive logging and error recovery mechanisms. Most importantly, this system needs to remain maintainable for years to come, surviving team changes and organizational growth.

Without proper project setup and governance from the beginning, even well-intentioned projects quickly run into serious problems. These aren't theoretical issues - they're pain points that plague countless ETL projects in production today. Let's examine what goes wrong and why establishing the right foundation from day one is critical to long-term success.
{% endhint %}

x

x

x

{% tabs %}
{% tab title="Common Pitfalls" %}
{% hint style="info" %}

#### Common Pitfalls That Derail Projects

The first major problem teams encounter is what's known as the "Works on My Machine" syndrome. This happens when developers hard-code paths directly into their transformations, such as writing `C:\Users\John\Documents\ETL\sales_data.csv` as the file location. Everything works perfectly on John's laptop, but when another developer checks out the code or when the job needs to run on a server, it fails immediately because that exact path doesn't exist. This problem multiplies across teams, with each developer maintaining their own slightly different version of the same transformations, making collaboration nearly impossible.

**Configuration chaos** represents another critical failure point. In many projects, database passwords end up stored directly in transformation files, visible to anyone who can access the code repository. Connection strings are copied and pasted into every job that needs database access, which means updating a password or server name requires hunting through potentially hundreds of files. When it's time to deploy to a new environment, teams realize they have production connection details scattered throughout development code, creating both security vulnerabilities and deployment nightmares. The result is either massive search-and-replace operations before each deployment or, worse, accidental connections to production databases from test environments.

**Version control** disasters occur when teams fail to properly integrate their ETL code with systems like Git. Without version control, there's no reliable backup of previous working versions. When something breaks, there's no way to compare the current code with what was working last week to identify what changed. Multiple developers end up overwriting each other's work, or competing versions of the same transformation exist in different locations with no clear indicator of which is the "real" one. Teams waste days recovering lost work or attempting to merge conflicting changes manually, and the project timeline slips while morale suffers.

**Deployment headaches** emerge when the only deployment process is manually copying files to production servers. Inevitably, someone forgets to copy a critical SQL script, or a configuration file gets left behind. Dependencies aren't documented, so the operations team doesn't know that the new customer dimension load requires a specific shell script or lookup table. Deployments that should take minutes stretch into hours of troubleshooting, often during planned downtime windows. Rollbacks become exercises in archaeology as teams try to remember exactly which files changed and which versions need to be restored.

Finally, **operational blind spots** leave teams flying blind in production. Jobs fail silently with no notifications sent to operations staff. When data quality issues are discovered weeks later, there's no audit trail showing what ran, when it ran, or why it failed. Root cause analysis becomes guesswork because there's no detailed logging of which records were processed, which were rejected, and what errors occurred. The business loses confidence in the data warehouse, and the technical team spends more time firefighting than building new features.
{% endhint %}

x
{% endtab %}

{% tab title="Key Concepts" %}
{% hint style="info" %}

#### Key Concepts

Understanding these four core concepts will transform how you build and maintain ETL solutions. Each concept addresses specific pain points we've just discussed while enabling professional-grade data integration practices.
{% endhint %}

{% hint style="info" %}
**Separation of Content and Configuration** is the principle of keeping your ETL logic completely separate from environment-specific settings. Your transformation should contain the data transformation rules, not hard-coded database connection strings or file paths. Instead, use variables like `${DB_CONNECTION}` or `${SOURCE_FILE_PATH}` throughout your code. These variables get their actual values from configuration files that exist outside your ETL code - one configuration file per environment.

When the same transformation runs in development, it reads the dev configuration file and connects to the development database at `localhost:3306/sales_dev`.

When that exact same transformation runs in production, it reads the production configuration file and connects to `prod-db-01:3306/sales`. You never modify the transformation itself when moving between environments.

This approach eliminates the "works on my machine" problem entirely, makes deployments straightforward, and allows you to update connection details in one place rather than hunting through hundreds of files.
{% endhint %}

{% hint style="info" %}
**Version Control Integration** means storing all your ETL artifacts- jobs, transformations, SQL scripts, documentation—in a Git repository just like software development teams do with application code. Every change is tracked with information about who made the change, when they made it, and why.

You can compare the current version of any transformation with previous versions to see exactly what changed. If a job that was working yesterday suddenly fails today, you can quickly identify that someone modified the customer validation logic this morning and review those specific changes.

Teams can work in parallel using feature branches without stepping on each other's toes, and you can tag specific versions as releases for deployment to production. When something breaks in production, you can roll back to the previous known-good version in minutes. This eliminates version control disasters and provides the foundation for professional deployment practices.
{% endhint %}

{% hint style="info" %}
The **Framework Pattern** provides a reusable layer of infrastructure code that wraps around your business logic. Think of it like the framework that web developers use - Django developers don't write user authentication code for every project; the framework provides it. Similarly, your ETL developers shouldn't write logging code, error handling, and job control logic in every transformation.

Instead, you build a framework once that provides these services, and all your ETL jobs use this framework. When a job runs, it doesn't execute directly; instead, a launcher job loads the configuration, initializes logging by creating a record in the job\_control table, executes your business logic, handles any errors that occur, updates the job completion status, and sends notifications if configured.

Your developers write only the transformations that implement business rules—loading customers, calculating metrics, or updating dimensions. The framework handles everything else. This means consistent behavior across all jobs, centralized improvements (fix a logging bug once and all jobs benefit), and developers who can focus on solving business problems rather than wrestling with infrastructure.
{% endhint %}

{% hint style="info" %}
**Job Restartability** enables failed jobs to resume from their point of failure rather than starting over from scratch. Consider a staging job that loads ten large tables sequentially. Everything runs perfectly through the first six tables - customer, product, store, time, promotion, and sales header have all loaded successfully, taking about thirty minutes total. Then the network hiccups, and the sales detail table fails after loading five million of ten million records. Without restartability, you fix the network issue and re-run the entire job, wasting thirty minutes reloading those six tables that already succeeded. With restartability, the framework tracks which steps completed successfully in the job\_control and step\_control tables. When you restart the job, it checks which tables already finished, skips them entirely, and resumes execution starting with the sales detail table that failed. The restart takes five minutes instead of thirty-five. This pattern saves enormous amounts of time when dealing with large data volumes, reduces the impact of transient failures like network timeouts, and makes your batch processing windows much more predictable. The step\_control table shows you exactly what succeeded and what failed, making troubleshooting straightforward rather than mysterious.
{% endhint %}

{% hint style="info" %}
These four concepts work together synergistically. Configuration separation enables the same code to run reliably in any environment. Version control tracks all changes and enables professional deployment practices. The framework provides consistent infrastructure services without cluttering business logic. Restartability minimizes the cost of failures, which inevitably occur in complex data integration scenarios. Together, they transform ETL development from an ad-hoc, error-prone process into a professional engineering discipline with predictable outcomes and maintainable solutions.
{% endhint %}
{% endtab %}
{% endtabs %}

x
{% endtab %}

{% tab title="2. Setup" %}

{% endtab %}

{% tab title="3. Configuration" %}

{% endtab %}

{% tab title="4. DI Framework" %}

{% endtab %}

{% tab title="5. Job Restartibility" %}

{% endtab %}

{% tab title="6. Monitoring & Logging" %}

{% endtab %}
{% endtabs %}

x

{% tabs %}
{% tab title="First Tab" %}
x
{% endtab %}

{% tab title="Second Tab" %}
x
{% endtab %}
{% endtabs %}

x


# Data Sources

Flat files, databases, storage, big data, and notebooks.

### Choose a data source

Use this page to orient yourself. Then jump into the specific connector docs:

{% hint style="info" %}

#### What “data source” means in PDI

In practice, a data source is either:

* A file format you parse (CSV, Excel, JSON, XML).
* A service you connect to (a DB, object store, cluster, or API).
  {% endhint %}

{% tabs %}
{% tab title="Flat Files" %}
{% hint style="info" %}

#### Flat files

Use flat files when your data arrives as CSV, TXT, fixed-width, JSON, or XML.

Start here: [Flat Files](/pentaho-data-integration/data-integration/data-sources/flat-files).
{% endhint %}

<figure><img src="/files/cbXkAVbVoLXxpwTDHqV6" alt=""><figcaption></figcaption></figure>

{% tabs %}
{% tab title="Structured" %}
{% hint style="info" %}

#### Structured

Structured data uses a predefined model. It is easy to validate and query.

Think tables, rows, and columns. Examples include SQL databases and well-formed CSV files.
{% endhint %}
{% endtab %}

{% tab title="Unstructured" %}
{% hint style="info" %}

#### Unstructured

Unstructured data has no consistent schema. Examples include PDFs, images, video, and free-form text.

You typically need parsing, extraction, or ML to use it.
{% endhint %}
{% endtab %}

{% tab title="Semi-structured" %}
{% hint style="info" %}

#### Semi-structured

Semi-structured data has a loose schema. It uses tags or keys to describe fields and hierarchy.

Common formats are JSON and XML.
{% endhint %}
{% endtab %}

{% tab title="Metadata" %}
{% hint style="info" %}

#### Metadata

Metadata is “data about data”. Examples include headers, schemas, and data dictionaries.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Databases" %}
{% hint style="info" %}

#### Databases

Pentaho connects to databases primarily through JDBC drivers. These drivers are the main interface for database communication.

Start here: [Databases](/pentaho-data-integration/data-integration/data-sources/databases).
{% endhint %}

<figure><img src="/files/IL7BZk89nZKZBXrAxInh" alt=""><figcaption><p>Database Connection</p></figcaption></figure>
{% endtab %}

{% tab title="Storage" %}
{% hint style="info" %}

#### Storage

Storage sources are cloud or network repositories. Examples include Amazon S3, Azure Blob Storage, and Google Cloud Storage.

In PDI, you typically connect through VFS. You can read and write across hybrid environments.

Start here: [Storage](/pentaho-data-integration/data-integration/data-sources/storage).
{% endhint %}

<figure><img src="/files/ygCVIRdwnpQ4roGXtxdP" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Big Data" %}
{% hint style="info" %}

#### Big Data

Big data sources require distributed compute. Common examples are Hadoop (HDFS, Hive, HBase), Spark, NoSQL, and Kafka.

PDI provides specialized steps and adapters for these platforms. This lets you transform data where it lives.

Start here: [Big Data](/pentaho-data-integration/data-integration/data-sources/big-data).
{% endhint %}

<figure><img src="/files/2uxLQWlognfwKxCe1BtQ" alt=""><figcaption><p>Types of Big Data</p></figcaption></figure>
{% endtab %}

{% tab title="Jupyter Notebook" %}
{% hint style="info" %}

#### Jupyter Notebook

Jupyter is a web-based notebook for code, visuals, and narrative text. It works well for exploratory analysis and prototyping.

In a PDI workflow, notebooks often handle advanced analysis. PDI handles production orchestration and scheduled pipelines.

Start here: [Jupyter Notebook](/pentaho-data-integration/data-integration/data-sources/jupyter-notebook).
{% endhint %}

<figure><img src="/files/WwNybMrPO7WNmGuxpry9" alt=""><figcaption><p>Jupyter Notebook</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Flat Files

How Data Integration handles Flat files ..

{% hint style="info" %}

#### **Flat Files**

**Structured flat files** are the most common type used in data integration, containing data organized in a consistent, predictable format with clearly defined fields and delimiters. Examples include **CSV (Comma-Separated Values)** files where each row represents a record and columns are separated by commas (e.g., `CustomerID,Name,Email,Purchase_Date`), **TSV (Tab-Separated Values)** files that use tabs as delimiters, and **fixed-width files** where each field occupies a specific number of characters (common in legacy mainframe systems). These files are ideal for Pentaho transformations because their predictable structure makes them easy to parse, with each row mapping directly to a database record and each column corresponding to a specific field.

**Unstructured flat files**, by contrast, contain free-form text without any predefined schema or organization, such as plain text documents, email bodies, or raw application log files that lack consistent formatting - these require more sophisticated text parsing and natural language processing techniques to extract meaningful data.
{% endhint %}

<figure><img src="/files/6fpACs9xGd9PiMzkxFmc" alt=""><figcaption><p>Flat Files</p></figcaption></figure>

{% hint style="info" %}
**Semi-structured flat files** occupy a middle ground, containing data with some organizational structure but without the rigid schema of databases or structured files. The most prominent examples are **JSON (JavaScript Object Notation)** files, which use key-value pairs and nested objects (e.g., `{"customer": {"id": 123, "orders": [{"item": "laptop", "price": 899}]}}`), and **XML (eXtensible Markup Language)** files that use hierarchical tags to define data relationships. These formats are self-describing and flexible, making them popular for APIs, web services, and modern application data exchange.

**Metadata in flat files** refers to descriptive information about the data itself - this can include header rows that define column names in CSV files, schema definitions that specify data types and constraints, file-level documentation about data source and creation date, or embedded comments that explain field meanings. In Pentaho, understanding and properly handling metadata is crucial for accurate data mapping, as it helps define how the ETL process should interpret field types (string vs. integer vs. date), handle null values, and validate data quality during transformation steps.
{% endhint %}


# Text

Ingesting Text Files ..

{% hint style="info" %}

#### Onboarding Text Files

**File Format and Encoding Issues** are among the most frequent obstacles. Inconsistent delimiters can occur when files mix separators (commas in some rows, tabs in others) or when text fields contain the delimiter character itself without proper quoting (e.g., "Smith, John" in a comma-delimited file). Line ending mismatches between Windows (`\r\n`), Unix/Linux (`\n`), and legacy Mac (`\r`) systems can cause Pentaho to misread record boundaries, resulting in merged or split rows. Character encoding problems arise when files created in different systems use incompatible encoding standards—a UTF-8 file with special characters (é, ñ, €) will display garbled text if read as ASCII, while files with emojis or international characters require UTF-8 or UTF-16 encoding to process correctly.

**Data Quality and Consistency Problems** can significantly impact ETL success. Missing values or incomplete records occur when rows have fewer fields than expected, forcing you to decide whether to skip these records, fill them with defaults, or flag them for review. Inconsistent data types within columns—such as a "Revenue" column containing both numeric values (`1250.50`) and text entries (`"N/A"` or `"pending"`)—will cause type conversion errors unless handled explicitly. Duplicate records, whether exact copies or near-duplicates differing only in whitespace or capitalization, require deduplication logic to maintain data integrity.

**Header Row and Structure Challenges** add another layer of complexity. Files may or may not include a header row defining column names, requiring you to manually map column positions to field names when headers are absent. Misaligned headers occur when the header row has a different number of fields than data rows (often due to merged cells in Excel exports or manual editing), causing field mapping errors. Dynamic file naming patterns with timestamps or sequential numbers (e.g., `sales_20241114_153045.csv` or `export_batch_0042.txt`) require wildcard matching or regular expressions in Pentaho to process files automatically without hardcoding specific filenames.

**Complex Data and Type Conversion Issues** require advanced handling techniques. Nested or hierarchical data embedded in flat files—such as pipe-delimited subcategories within a comma-delimited file (`Product,Categories\nLaptop,"Electronics|Computers|Hardware"`)—needs parsing logic to extract and normalize. Multi-line records where a single logical record spans multiple physical lines (common in address fields or comments) must be reassembled before processing. Date and time format inconsistencies are particularly troublesome, as files may contain dates in various formats (`MM/DD/YYYY`, `DD-MM-YYYY`, `YYYY-MM-DD`) or ambiguous formats (`01/02/2024` could be January 2nd or February 1st depending on regional settings), requiring explicit date parsing with format masks in Pentaho to avoid misinterpretation.
{% endhint %}

***


# Text File Input

Ingest semi-structured text files into clean rows.

{% hint style="warning" %}

#### Workshop - Text File Input

Real-world data rarely arrives in perfect, structured formats ready for database loading. Organizations frequently receive orders, invoices, and other business documents as unstructured or semi-structured text files that require significant transformation before they can be analyzed. Learning to parse, cleanse, and structure these files is an essential skill for any data integration professional.

In this hands-on workshop, you'll work with Steel Wheels' order data delivered in a challenging text format. You'll build a complete transformation pipeline that takes messy, multi-line text records and converts them into clean, structured rows suitable for database insertion. This workshop introduces several powerful PDI steps for text manipulation, pattern matching, and data formatting techniques you'll use repeatedly when integrating data from legacy systems, EDI feeds, or flat file exports.

**What you'll do**

* Configure the Text File Input step to read unstructured text data
* Use the Flattener step to convert multi-line records into single rows
* Apply Regular Expressions (RegEx) to extract specific data patterns and create capture groups
* Implement the Replace in String step to remove unwanted text and formatting
* Perform explicit data type conversions using the Select Values step
* Format currency values and dates for proper database storage
* Build a complete text processing pipeline from raw input to structured output

By the end of this workshop, you'll understand the multi-step process required to onboard flat files into database tables. You'll have practical experience with pattern matching, string manipulation, and data type conversion - core competencies that enable you to tackle even the most challenging text file formats. Instead of relying on manual data clean-up or complex pre-processing scripts, you'll build automated, repeatable transformations that handle messy data with confidence.

**Prerequisites:** Understanding of basic transformation concepts (steps, hops, preview); Pentaho Data Integration installed and configured

**Estimated time:** 30 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/6b3348c091764d08806280206bd53434?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Text File Input
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/84xel1qQzg06jp8Mh91u" %}

{% file src="/files/UqZMjNp6V6BvE8WJ2NhP" %}

***

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

<figure><img src="/files/F1GRVzsGoAUBIeBkuqQ6" alt=""><figcaption><p>Text files</p></figcaption></figure>

***

{% hint style="info" %}
Review the input file first. It will guide your parsing approach.
{% endhint %}

<figure><img src="/files/DOHzBoPuYBiDXTdSsz5A" alt="" width="563"><figcaption><p>orders.txt</p></figcaption></figure>

{% hint style="info" %}
What to notice:

* Each order spans multiple lines.
* Line 3 contains two values: order status and order date.
* Order value includes a currency symbol ($).
* There is inconsistent whitespace.
  {% endhint %}

{% hint style="info" %}
**Approach**

You will:

* Flatten multi-line records into a single row.
* Extract values into new fields (capture groups).
* Clean strings (remove labels and currency symbols).
* Set data types and formats (date and number).
  {% endhint %}

{% embed url="<https://www.loom.com/share/0646ed96c4734d16b6721c902f9c7e0e?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
String Operations
{% endembed %}

***

{% tabs %}
{% tab title="1. Text File Input" %}
{% hint style="info" %}

#### Text File Input

Use **Text file input** to read the raw lines from `orders.txt`. Treat each line as a single string field for now.
{% endhint %}

1. Start Pentaho Data Integration (Spoon).

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. In the **Design** tab, expand the `Input` category.
3. Drag **Text file input** onto the canvas.

{% hint style="info" %}
Tip: You can also search for `Text file input`.
{% endhint %}

4. Double-click the step. Configure the file path:

<figure><img src="/files/9vUudlE2Z5DWX4VxcjVs" alt=""><figcaption><p>Add path to file</p></figcaption></figure>

{% hint style="info" %}
Because the sample file is located in the same directory where the transformation resides, a good approach to naming the file in a way that is location independent is to use a system variable to parameterize the directory name where the file is located. In our case, the complete filename is:

`${Internal.Transformation.Filename.Directory}/orders.txt`
{% endhint %}

5. Select the **Content** tab. Configure it like this:

<figure><img src="/files/RA7fuTe1S2h6ua5VNETJ" alt="" width="563"><figcaption><p>Text file input - Content</p></figcaption></figure>

6. Select **Fields**. Select **Get Fields**.

<figure><img src="/files/TxKfbj7wNBNdeDGZ9Jge" alt="" width="563"><figcaption><p>Text File input - Fields</p></figcaption></figure>

{% hint style="info" %}
The step returns one field named `Field1`. It has type **String**.
{% endhint %}

7. Optional: rename the step to **Read orders**.
8. Select **OK**.
   {% endtab %}

{% tab title="2. Row flattener" %}
{% hint style="info" %}

#### Row Flattener

Use **Flattener** to turn repeating lines into a single output row.
{% endhint %}

1. Drag **Flattener** onto the canvas.
2. Create a hop from **Read orders**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/BG9iExrK49qyceewag7o" alt="" width="375"><figcaption><p>Row flattener</p></figcaption></figure>

4. Optional: rename the step to **Flatten rows**.
5. Select **OK**.

{% hint style="info" %}
The data has now been flattened into records. This step enables you to define new target fields that match the number of repeating records. So Target field 1 will map to repeating record 1, and so on.
{% endhint %}
{% endtab %}

{% tab title="3. RegEx Evaluation" %}
{% hint style="info" %}

#### RegEx Evaluation

This step type allows you to match the String value of an input field against a text pattern defined by a regular expression. Optionally, you can use the regular expression step to extract substrings from the input text field matching a portion of the text pattern into new output fields. This is known as "capturing".

In our example, we’re going to extract and create two capture groups order\_status and order\_date based on the regex expression: (Delivered|Returned):(.+)
{% endhint %}

1. Drag **RegEx Evaluation** onto the canvas.
2. Create a hop from **Flatten rows**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/YKZxq2ElQIhZewVdThy0" alt=""><figcaption><p>RegEx Evaluation</p></figcaption></figure>

{% hint style="warning" %}
Set **Trim** to **both** for each field. This removes leading and trailing whitespace.
{% endhint %}

4. Optional: rename the step to **Parse status and date**.
5. Select **OK**.

***

{% hint style="info" %}
**Summary**

* This RegEx uses 2 constructs, denoted by the brackets, and separated by a full colon.
* `(Delivered|Returned)` matches either status.
* `(.+)` matches any character sequence.
* Use **Test RegEx** to verify capture groups.
  {% endhint %}

A good introduction can be found at:

{% embed url="<https://regex101.com/>" %}
Link to: Online RegEx engine
{% endembed %}
{% endtab %}

{% tab title="4. Replace in String" %}
{% hint style="info" %}

#### Replace in string

Replace in string is a simple search and replace. It also supports regular expressions and group references. Group references are picked up in the replace by string as $n where n is the number of the group.

Time to tidy up the order\_value stream field data. In this step, you replace the `Order Value:` label with an empty string.
{% endhint %}

1. Drag **Replace in string** onto the canvas.
2. Create a hop from **Parse status and date**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/PzMW5SIJJiKw0a406aOt" alt=""><figcaption><p>Replace in String</p></figcaption></figure>

4. Optional: rename the step to **Clean order value**.
5. Select **OK**.

{% hint style="danger" %}
Use the exact label text, including the trailing space. If you enable regular expressions, use `Order Value:\s*`.
{% endhint %}
{% endtab %}

{% tab title="5. Select Values" %}
{% hint style="info" %}

#### Select values

The Select Values step is useful for selecting, removing, renaming, changing data types and configuring the length and precision of the fields on the stream. These operations are organized into different categories:

* Select and Alter — Specify the exact order and name in which the fields should be placed in the output rows
* Remove — Specify the fields that should be removed from the output rows
* Metadata — Change the name, type, length, and precision (the metadata) of one or more fields
  {% endhint %}

1. Drag **Select values** onto the canvas.
2. Create a hop from **Clean order value**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/5oXib9jaud7GNDsvLL7T" alt=""><figcaption><p>Select values</p></figcaption></figure>

| Fieldname    | Data Type | Format   |
| ------------ | --------- | -------- |
| order\_value | Number    | #.00     |
| order\_date  | Date      | MMM yyyy |

{% hint style="info" %}
Optional: rename the step to **Set data types**.
{% endhint %}
{% endtab %}

{% tab title="6. RUN" %}
{% hint style="info" %}

#### Run the transformation

Run the transformation locally.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Select the **Preview data** tab.

<figure><img src="/files/xcfKSWwtDM6kgU2OT6l7" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="success" %}
You should now have clean fields such as `order_date`, `order_status`, and `order_value`.
{% endhint %}
{% endtab %}
{% endtabs %}


# Text File Output

Output text files ..

{% hint style="warning" %}

#### Workshop - Text File Output

Reading files is only half the job. You also need to generate files for users and systems.

In this workshop, you build a transformation that writes a customer survey for Steel Wheels. You will build the survey from multiple streams. You will parameterize it with a runtime customer name.

**What you'll do**

* Read a customer name from a transformation argument
* Build header and body sections with static rows and file-driven rows
* Format text using User Defined Java Expression
* Merge streams in a predictable order with Append streams
* Write the final output with Text file output

**Prerequisites:** Understanding of basic transformation concepts (steps, hops, preview). Complete [Text File Input](/pentaho-data-integration/data-integration/data-sources/flat-files/text/text-file-input) first.

**Estimated time:** 35 minutes
{% endhint %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/9ZFuevKDcD2MR3sJhbxr" %}

{% file src="/files/j3nocWDNeiD3pk9nBZlw" %}

***

<figure><img src="/files/i4IlMya45cJuiPBpFo5l" alt=""><figcaption><p>Survey - Text File Output</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Get System Info" %}
{% hint style="info" %}

#### Get System Info

Use **Get System Info** to read a runtime argument. We will treat the argument as the customer name.
{% endhint %}

{% embed url="<https://www.loom.com/share/de06920fdcd84bc2b1fe63454afc8df8?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Get System Info
{% endembed %}

1. Start Pentaho Data Integration (Spoon).

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag **Get System Info** onto the canvas.
3. Double-click the step. Configure it like this:

<figure><img src="/files/FSj8ure1CbRZDnLMMUTQ" alt="" width="375"><figcaption><p>Get system info</p></figcaption></figure>

4. Select **OK**.
   {% endtab %}

{% tab title="2. User Defined Java Expression" %}
{% hint style="info" %}

#### User Defined Java Expression

Use **User Defined Java Expression** to format the header line. It will combine a label with the customer name.
{% endhint %}

1. Drag **User Defined Java Expression** onto the canvas.
2. Create a hop from **Get System Info**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/WkYXRwo6AzjNRBtraPiH" alt=""><figcaption><p>UDJE</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
This replaces the original argument value with formatted text. The output field is still named `text`.
{% endhint %}

{% embed url="<https://docs.oracle.com/javase/tutorial/java/nutsandbolts/opsummary.html>" %}
Link to Operators
{% endembed %}
{% endtab %}

{% tab title="3. Data Grid" %}
{% hint style="info" %}

#### Data Grid

Use **Data Grid** to add static survey header lines. This keeps the top-of-file content inside the transformation.

Configure the field metadata on **Meta**. Enter the rows on **Data**.
{% endhint %}

1. Drag **Data Grid** onto the canvas.
2. Double-click the step. Configure it like this:

<div><figure><img src="/files/OWgpmwx6tpJcGWkePViX" alt=""><figcaption><p>Data grid - text</p></figcaption></figure> <figure><img src="/files/0X2PnoHrYBFHB3p8U7cK" alt=""><figcaption><p>Data grid - data</p></figcaption></figure></div>

3. Select **OK**.
   {% endtab %}

{% tab title="4. Append streams (head)" %}
{% hint style="info" %}

#### Append streams

Use **Append streams** when order matters. It outputs all rows from the first hop. It then outputs all rows from the second hop.

Both input streams must have the same field names and types.
{% endhint %}

1. Drag **Append streams** onto the canvas.
2. Create hops from **User Defined Java Expression** and **Data Grid**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/pjbrLQgWBddoqKbZj6kR" alt="" width="375"><figcaption><p>Append</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
To append streams, keep the layout consistent. In this workshop, every stream uses a single `text` field.
{% endhint %}

{% hint style="info" %}
If order does not matter, use a step that performs a union of streams instead.
{% endhint %}

{% hint style="warning" %}
Make sure the header stream is the **first** input hop. Append streams will output that stream first.
{% endhint %}
{% endtab %}

{% tab title="5. Text file input (questions)" %}
{% hint style="info" %}

#### Text file input

Use **Text file input** to read the question list from a file. Each question becomes one row.
{% endhint %}

1. Drag **Text file input** onto the canvas.
2. Double-click the step. Configure it like this:

<figure><img src="/files/8EkbmGtGwvdMHJFpy3e1" alt=""><figcaption><p>Text file input</p></figcaption></figure>

File: `${Internal.Transformation.Filename.Directory}/questions.txt`

<figure><img src="/files/9ZIxp4IP0dF2wEo8LaIn" alt=""><figcaption><p>Text file input - Content</p></figcaption></figure>

{% hint style="info" %}
Use a **Tab** delimiter. Enable **row numbers** to generate question numbers.
{% endhint %}

3. On **Fields**, rename the output field to `text`:

<figure><img src="/files/vawmgI961L4E48Zs5T2l" alt=""><figcaption><p>Text file input - Fields</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
Each row now contains a question in `text`. The row number field (for example `question_num`) identifies the question number.
{% endhint %}
{% endtab %}

{% tab title="6. User Defined Java Expression (number questions)" %}
{% hint style="info" %}

#### User Defined Java Expression

Use a second **User Defined Java Expression** to prefix each question line with its question number.
{% endhint %}

1. Drag **User Defined Java Expression** onto the canvas.
2. Create a hop from **Text file input (questions)**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/jDqtRbvQod5yrnSma2Ke" alt=""><figcaption><p>UDJE - concat question numbers</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
This overwrites `text` with a numbered question like `1. How did we do?`.
{% endhint %}
{% endtab %}

{% tab title="7. Select values" %}
{% hint style="info" %}

#### Select values

The Select Values step is useful for selecting, removing, renaming, changing data types and configuring the length and precision of the fields on the stream.

These operations are organized into different categories:

* Select and Alter — Specify the exact order and name in which the fields should be placed in the output rows
* Remove — Specify the fields that should be removed from the output rows
* Meta-data — Change the name, type, length, and precision (the metadata) of one or more fields
  {% endhint %}

1. Drag **Select values** onto the canvas.
2. Create a hop from **User Defined Java Expression (number questions)**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/T63zgCxYo52BzqgqcsGd" alt="" width="375"><figcaption><p>Select values - remove question_num</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
Remove the question number field so both streams have the same layout. You need a single `text` field before you append.
{% endhint %}
{% endtab %}

{% tab title="8. Append streams (body)" %}
{% hint style="info" %}

#### Append streams

Append the survey header stream to the numbered question stream.
{% endhint %}

1. Drag **Append streams** onto the canvas.
2. Create hops from **Append streams (head)** and **Select values**.
3. Double-click the step. Configure it like this:

<figure><img src="/files/sdWc8d0WPSSEnEcma0s5" alt="" width="375"><figcaption><p>Append</p></figcaption></figure>

4. Select **OK**.

{% hint style="info" %}
You now have one stream. It contains one field named `text`.
{% endhint %}
{% endtab %}

{% tab title="9. Text file output" %}
{% hint style="info" %}

#### Text file output

Use **Text file output** to write the survey file to disk.
{% endhint %}

{% hint style="warning" %}
Do not run multiple copies of this step against the same output file. Use **Include stepnr in filename** if you need parallel output.
{% endhint %}

1. Drag **Text file output** onto the canvas.
2. Create a hop from **Append streams (body)**.
3. Double-click the step. Set the file name:

`Filename: ${Internal.Transformation.Filename.Directory}/survey`

{% hint style="info" %}
Set **Extension** to `txt` if your output should be `survey.txt`.
{% endhint %}

<figure><img src="/files/5O4spQtjpawOuSHJA1HT" alt="" width="563"><figcaption><p>Text file output - Content</p></figcaption></figure>

4. On **Fields**, select **Get Fields**.

<figure><img src="/files/iUwRGcBtr2tpxrdA9O9i" alt="" width="563"><figcaption><p>Text file output - Fields</p></figcaption></figure>

5. Select **OK**.
   {% endtab %}

{% tab title="10. RUN" %}
{% hint style="info" %}

#### Run the transformation

Run the transformation locally. Pass a customer name as an argument.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. Open **Arguments (legacy)**. Enter a customer name.

<figure><img src="/files/6sWFwTZwEy8zGzos7lQD" alt=""><figcaption><p>Enter argument</p></figcaption></figure>

3. Select **Run**.
4. Open the **Preview data** tab.

<figure><img src="/files/vAsYgwfqBcic18tOlgC1" alt=""><figcaption><p>Preview data</p></figcaption></figure>

5. Open the generated survey file in your transformation folder.

{% hint style="info" %}
This workshop reinforces the rule for merging streams:

* Keep the same field layout (names and order).
* Keep matching data types.
  {% endhint %}
  {% endtab %}
  {% endtabs %}


# Excel

Time for some smoke & mirrors ..

{% hint style="info" %}

#### **Microsoft Excel**

**Data Input and Reading** is handled through Pentaho's dedicated Excel input steps. The **"Excel Input"** and **"Microsoft Excel Input"** steps allow you to read both legacy `.xls` (Excel 97-2003) and modern `.xlsx` (Excel 2007+) formats directly into your transformations. You can specify which worksheets to read by name or index, define specific cell ranges (e.g., `A1:G500` or `Sheet2!B5:F100`), skip header rows, and process multiple sheets within a single workbook either sequentially or in parallel. These steps automatically detect column names from the first row when configured and can handle complex workbook structures including merged cells, hidden sheets, and password-protected files (when credentials are provided).

**Data Output and Writing** is accomplished using the **"Excel Output"** and **"Microsoft Excel Writer"** steps, which support creating brand-new Excel files or appending data to existing workbooks. You have fine-grained control over formatting options including sheet names, cell styles, column widths, header row formatting, and data type preservation (dates, numbers, text). The output steps can dynamically name files using variables, write multiple sheets to a single workbook, and handle large datasets by streaming data rather than loading everything into memory - though extremely large files (100,000+ rows) may benefit from CSV export instead due to Excel's file size and performance limitations.

**Template-Based Reporting and Dynamic Generation** is a powerful feature where pre-formatted Excel templates serve as the foundation for automated report generation. You can create polished templates with company branding, charts, pivot tables, and conditional formatting, then use PDI to inject live data into specific cells or named ranges while preserving all formatting and formulas. This approach is ideal for recurring business reports (monthly sales summaries, financial dashboards, operational KPIs) where the layout remains constant but the data refreshes from your data warehouse or transactional systems. The **"Excel Writer"** step specifically supports template injection, allowing you to maintain separate template files managed by business users while PDI handles the data population.

**Data Transformation and Business Logic Migration** enables you to replicate Excel-based calculations within your ETL processes. Complex formulas used in Excel spreadsheets - such as `VLOOKUP`, `SUMIF`, nested `IF` statements, or custom business calculations - can be translated into PDI's **Calculator**, **Formula**, or **User Defined Java Expression** steps, moving the logic from desktop spreadsheets into enterprise-grade, version-controlled, and auditable ETL workflows. This migration is crucial when organizations need to scale beyond Excel's limitations, ensure data governance, or eliminate "spreadsheet hell" where critical business processes depend on error-prone manual spreadsheet manipulation.

**Metadata Extraction and Schema Discovery** capabilities allow PDI to intelligently inspect Excel files before processing. The input steps can automatically detect sheet names, column headers, data types (text, numeric, date, boolean), and even sample data to help you configure field mappings. This is particularly useful when dealing with dynamically structured Excel files where the schema isn't known in advance, or when building generic transformations that need to adapt to different Excel file layouts from various departments or external partners.

**Error Handling and Data Quality Management** provides several mechanisms to deal with Excel-specific challenges. You can configure how to handle empty cells (treat as null vs. empty string vs. zero), what to do with formula errors (`#DIV/0!`, `#VALUE!`, `#REF!`—either skip the row, substitute a default value, or log the error), and how to process cells with data validation errors or out-of-range dates. The **"Filter Rows"** and **"Switch/Case"** steps can route problematic records to error handling streams, while the **"Write to Log"** step helps identify which specific rows or cells caused issues, making it easier to work with business users to correct source data quality problems in their Excel files.
{% endhint %}


# Excel Writer

Working with Excel ..

{% hint style="warning" %}

#### Workshop - Excel Writer

Excel reports often need templates, charts, and fixed layouts.

Here you populate a pre-formatted Sales and Expenses report. You will write multiple sections into one workbook. You will control execution order so writers do not conflict.

**What you'll do**

* Use a template workbook and write to fixed cell positions
* Write a report header with **Generate rows**
* Read sales and expense rows from text files
* Block parallel flows before writing to the same file
* Write multiple sections with **Microsoft Excel Writer**

By the end, you will know how to write into an Excel template safely. You will also know when to block parallel flows.

**Prerequisites:** Understanding of basic transformation concepts (steps, hops, preview). Complete [Text File Input](/pentaho-data-integration/data-integration/data-sources/flat-files/text/text-file-input) first.

**Estimated time:** 35 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/770dd4bae75049eeab247bec2bf6fbcc?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Advanced Excel Writer scenario
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/Jlm87v9GAlcd391Rmc0x" %}

{% file src="/files/anG9Vsj7bhfF7mW9cF3M" %}

{% file src="/files/PknQoJlSxgCKMMBUKxBk" %}

{% file src="/files/53NBz48h0Fad63DEj9sf" %}

***

<figure><img src="/files/FSJJwDhqyLt6uLAXZhtT" alt="" width="375"><figcaption><p>Sales &#x26; Expenses</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Excel Template" %}
{% hint style="info" %}

#### Excel template

The various stages of the transformation write data to a template.xlsx. The template has 2 worksheets:

* Sales Chart - this worksheet creates a 3D stacked graph
* SourceData - worksheet
  {% endhint %}

1. Open `template.xlsx` in Excel:

<figure><img src="/files/Ti55UrSktE7weQm7ML2S" alt=""><figcaption><p>Blank template</p></figcaption></figure>

{% hint style="info" %}
SourceData - the datasheet. Transformations write to the required cells that are used to create the graph.
{% endhint %}

<figure><img src="/files/IxLFIbezFZ7M5xwLKPE8" alt=""><figcaption><p>Blank SourceData</p></figcaption></figure>
{% endtab %}

{% tab title="2. Write Year" %}
{% hint style="info" %}

#### Write year

The first workflow is to write the current Year to the SourceData worksheet in the template.xlsx

You can change the year value.
{% endhint %}

<figure><img src="/files/wGgwSVZf8yAU8tg3DUNX" alt="" width="375"><figcaption><p>Year</p></figcaption></figure>

{% tabs %}
{% tab title="1. Generate Rows - Year" %}
{% hint style="info" %}

#### Generate rows - year

Generate rows outputs a fixed number of rows. Here you output a single row that contains the report year.
{% endhint %}

1. Start Pentaho Data Integration (Spoon).

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag the ‘Generate Rows’ step onto the canvas.
3. Double-click on the step, and configure the following properties:

<figure><img src="/files/LHDK2Do8RUGvKsTAGGJF" alt=""><figcaption><p>Generate rows - Year</p></figcaption></figure>

4. Close Step.

***

{% hint style="info" %}
**Summary**

* Generates a record that holds the Year value – 2023 – in the year stream field.
* The Excel template will also need to be formatted yyyy to interpret the Date.
  {% endhint %}
  {% endtab %}

{% tab title="2. Excel Writer - Year" %}
{% hint style="info" %}

#### Excel Writer - year

Microsoft Excel Writer writes incoming rows into an Excel workbook. Use `xlsx` when you work with templates and charts.
{% endhint %}

1. Drag the ‘Excel writer’ step onto the canvas.
2. Create a hop from the ‘Year’ step.
3. Double-click on the step, and configure the following properties:

<figure><img src="/files/3qwOdohUXR707dfAgbAg" alt=""><figcaption><p>Excel writer - Year</p></figcaption></figure>

{% hint style="info" %}
Use these paths:

* Output: `${Internal.Transformation.Filename.Directory}/Sales_and_Expenses_2023.xlsx`
* Template: `${Internal.Transformation.Filename.Directory}/template.xlsx`

Select **Replace with new output file** while you develop. It resets the workbook on every run.
{% endhint %}

4\. Click on the Content tab, and configure the following properties:

<figure><img src="/files/IuCQCssmqSMz2doCQ4Kn" alt=""><figcaption><p>Excel writer - cell</p></figcaption></figure>

5. Click on ‘Get Fields’ button.
6. Click OK.
   {% endtab %}
   {% endtabs %}
   {% endtab %}

{% tab title="3. Write Sales" %}
{% hint style="info" %}

#### Write sales

Write sales rows into the same workbook. Use a blocking step so the year write completes first.
{% endhint %}

<figure><img src="/files/Cqt2UNbfSs6Gg6gkVI5H" alt="" width="375"><figcaption><p>Write Sales</p></figcaption></figure>

{% tabs %}
{% tab title="1. Text File Input - Read Sales" %}
{% hint style="info" %}

#### Text file input - read sales

Read the sales dataset from the workshop file. Keep the header row enabled so field names match the template.
{% endhint %}

1. Drag the ‘Text file input’ step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/BqY2Tfr8W1XWGxU0ZI8j" alt=""><figcaption><p>Text file input - sales</p></figcaption></figure>

3. Click on the Content tab, and configure the following properties:

<figure><img src="/files/BasBoqI1JvfSM8uD4TXL" alt=""><figcaption><p>Text file input - Content</p></figcaption></figure>

{% hint style="info" %}

* Ensure the Header is selected.
* No empty rows
* Mixed Format
  {% endhint %}

4. Click on the Fields tab, and click on ‘Get Fields’ button:

{% hint style="info" %}
Returns the Header values as stream fields.
{% endhint %}

5. Click OK.
   {% endtab %}

{% tab title="2. Block until Step Finish - Wait Year" %}
{% hint style="info" %}

#### Block until steps finish - wait year

This step waits for specific steps to finish. Use it to prevent parallel writers.
{% endhint %}

1. Drag the ‘Block this step until steps finish’ step onto the canvas.
2. Create a hop from the ‘Read Sales’ step.
3. Double-click on the step, and configure the following properties:
   * Watch step: the step that writes the year (copy `0`)
   * If you used **Get steps**, remove everything except the year writer step

<figure><img src="/files/qGk6K5fcKfHqPPkIY1g3" alt=""><figcaption><p>Block step</p></figcaption></figure>

{% hint style="info" %}
This will result in the workflow being blocked until the Write Year step has been completed.
{% endhint %}
{% endtab %}

{% tab title="3. Excel Writer - Write Sales" %}
{% hint style="info" %}

#### Excel Writer - write sales

Write sales rows into the existing workbook. Use **Use existing file for writing**.
{% endhint %}

1. Drag the ‘Excel writer’ step onto the canvas.
2. Create a hop from the ‘Wait Year’ step.
3. Double-click on the step, and configure the following properties:

<figure><img src="/files/cVdU34EvufQXNEDsZfIs" alt=""><figcaption><p>Excel writer - Sales</p></figcaption></figure>

{% hint style="info" %}
Use the same output path you used in the year writer:

* Output: `${Internal.Transformation.Filename.Directory}/Sales_and_Expenses_2023.xlsx`

Select **Use existing file for writing**.
{% endhint %}

4. Click on the Content tab, and configure the following properties:

<figure><img src="/files/w9jmyBSeSuPrgeDEXq3x" alt=""><figcaption><p>Excel writer - Content</p></figcaption></figure>

5. Click on the ‘Get Fields’ button.
6. Delete the productline field, as its not required. The template already has the fieldname and you are just writing the data, starting at cell B5.
7. Click OK.
   {% endtab %}
   {% endtabs %}
   {% endtab %}

{% tab title="4. Write Expenses" %}
{% hint style="info" %}

#### Write expenses

Write expense rows into the same workbook. Block until the sales write completes.
{% endhint %}

<figure><img src="/files/2AQRU355TbuKL0VxaHMF" alt="" width="375"><figcaption><p>Write Expenses</p></figcaption></figure>

{% tabs %}
{% tab title="1.  Text File Input - Read Expenses" %}
{% hint style="info" %}

#### Text file input - read expenses

Start with loading the Expenses data into the data stream.
{% endhint %}

<figure><img src="/files/K2qz8otBwaANaSQ6GFe0" alt=""><figcaption><p>Text File input - Expenses</p></figcaption></figure>

<figure><img src="/files/cMFYw4SWBB9MbosMGVwv" alt=""><figcaption><p>Text File input - Content</p></figcaption></figure>

<figure><img src="/files/q0FBVOvulVBfMlGRbzot" alt=""><figcaption><p>Text File input - Fields</p></figcaption></figure>
{% endtab %}

{% tab title="2. Block until Step finish - Sales" %}
{% hint style="info" %}

#### Block until steps finish - wait sales

Wait until sales data has finished writing to the workbook.
{% endhint %}

<figure><img src="/files/JhxMS8lIyGKqSjMwgGVU" alt=""><figcaption><p>Block 'Write Sales'</p></figcaption></figure>
{% endtab %}

{% tab title="3. Excel Writer - Expenses" %}
{% hint style="info" %}

#### Excel Writer - Expenses

Write expenses to `SourceData`.
{% endhint %}

<figure><img src="/files/5E3jnZFx6A8mtczRvzEf" alt=""><figcaption><p>Excel writer - Expenses</p></figcaption></figure>

<figure><img src="/files/2Ld6gCeAmNbZi2O3uRqm" alt=""><figcaption><p>Excel writer - Content</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="5. RUN" %}
{% hint style="info" %}

#### Run the transformation

Steps initialize in parallel. Use blocking steps to prevent concurrent writes to the workbook.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Open the Sales\_and\_Expenses\_2023.xlsx file.

<figure><img src="/files/mN1bDhC0kJ0nqssVeJ7F" alt=""><figcaption><p>Excel Book</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# XML

Data exchange & storage ..

{% hint style="info" %}

#### **XML & XPath**

XML (eXtensible Markup Language) is a versatile format for structuring and storing data, widely used in various applications and data exchange scenarios. One of the powerful tools for working with XML is XPath (XML Path Language), which provides a way to navigate and extract specific data from XML documents.

XPath is a set of rules used for getting information from an XML document. In XPath, XML documents are treated as trees of nodes. There are several types of nodes; elements, attributes, and texts are some of them. As an example, document, and order are some of the nodes in the sample file.

Among the nodes there are relationships. A node has a parent, zero or more children, siblings, ancestors, and descendants depending on where the other nodes are in the hierarchy. To select a node in an XML document, you should use a path expression relative to a current node.
{% endhint %}

<figure><img src="/files/kdosGADZDEXvA4woBIto" alt="" width="375"><figcaption><p>X path</p></figcaption></figure>

{% embed url="<https://www.w3schools.com/xml/xpath_intro.asp>" %}
Link to Xpath tutorial
{% endembed %}

***


# Read XML

XML data sources ..

{% hint style="warning" %}

#### Workshop - Read XML

Read XML from a file, a URL, or a field value. Use **Get data from XML**.

**What you’ll do**

* Read XML from a local file.
* Read XML from a URL (URI).
* Use **XPath** to select nodes and fields.
* Use **Get Fields** to infer the XML structure.
* Debug a data type mismatch using the logs.

**Prerequisites:** Basic transformations. Basic XML (elements, attributes, hierarchy). PDI installed.

**Estimated time:** 30 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/85ad9973848041b9b8447ed45cffc09c?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Get Data from XML
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download these files before you start:

* The sample XML input.
* The starter transformation (optional).
  {% endhint %}

{% file src="/files/cXoi8spmhLLXNFUMxNWt" %}

***

<figure><img src="/files/Df8l3Bg4KIq875nD4ReF" alt="" width="375"><figcaption><p>Read XML</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. XML - File" %}
{% hint style="info" %}

#### **XML - File**

In this workflow, an XML **file** is parsed via **XPath** to retrieve the dataset.
{% endhint %}

<figure><img src="/files/Nfy0uxdm87mjDQK7wmbL" alt="" width="375"><figcaption><p>Get data from XML - file</p></figcaption></figure>

<figure><img src="/files/ep2niXC1hpTfWwHQ9pHu" alt="" width="375"><figcaption><p>orders.xml</p></figcaption></figure>

{% tabs %}
{% tab title="1. Get data from XML" %}
{% hint style="info" %}

#### **Get data from XML**

This step provides the ability to read data from any type of XML file using XPath specifications.
{% endhint %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag the ‘Get data from XML’ step onto the canvas.
3. Double-click on the step, and configure the following properties:

<figure><img src="/files/3Rwfg6wwzMnvuF52ohRC" alt=""><figcaption><p>XML - file</p></figcaption></figure>

4. Click on the Content tab, and configure the following properties:

<figure><img src="/files/0jLOQx1ereeA6tCTCoiZ" alt=""><figcaption><p>XPath</p></figcaption></figure>

5. Click on the Fields tab, and then on the ‘Get Fields’ button.

<figure><img src="/files/Z6I0XT7n3TC8bdlswFUa" alt=""><figcaption><p>XML - fields</p></figcaption></figure>

6. Click OK.
   {% endtab %}

{% tab title="2. Dummy" %}
{% hint style="info" %}

#### **Dummy**

The Dummy step does not do anything. Its primary function is to be a placeholder for testing purposes. For example, to have a transformation, you need at least two steps connected to each other.
{% endhint %}

1. Drag a ‘Dummy’ step onto the canvas.
2. Create a hop from the ‘Get data from XML’ step.
3. Close the Step.
   {% endtab %}

{% tab title="3. RUN" %}
{% hint style="info" %}

#### **RUN Transformation**

The workshop illustrates how to ingest an XML data source. The XML can either stream from:

* a previous step (typically a URL)
* a file
* a stream field (XML stored in a field)
  {% endhint %}

{% hint style="warning" %}
Remember to disable the hops on the second workflow.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Preview the data.

<figure><img src="/files/aiBBvxgp1W4V11uFDJEu" alt=""><figcaption><p>Preview data</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. XML - URI" %}
{% hint style="info" %}
In this workflow, a **URL** to an XML data source is parsed via **XPath** to retrieve the dataset.
{% endhint %}

<figure><img src="/files/0hZYvUgGKOiXLX0oLR70" alt="" width="375"><figcaption><p>Get XML - URL</p></figcaption></figure>

<figure><img src="/files/tSgfm42MAaiGb98J7Uc9" alt="" width="375"><figcaption><p><a href="https://www.w3schools.com/xml/plant_catalog.xml">https://www.w3schools.com/xml/plant_catalog.xml</a></p></figcaption></figure>

{% tabs %}
{% tab title="1. Generate rows - Pass URL" %}
{% hint style="warning" %}
In this workshop, you pass the URL in a data stream field.

Copy the URL to your clipboard. You will paste it into the XPath dialog.
{% endhint %}

{% hint style="info" %}

#### Generate rows

Generate rows outputs a specified number of rows. By default, the rows are empty; however, they can contain several static fields. This step is used primarily for testing purposes. It may be useful for generating a fixed number of rows, for example, you want exactly 12 rows for 12 months.
{% endhint %}

1. Drag the ‘Generate Rows’ step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/H5Lp8FsHWG5llczDW4OE" alt=""><figcaption><p>Pass URL in data stream field</p></figcaption></figure>
{% endtab %}

{% tab title="2. Get data from XML - Read URL" %}
{% hint style="info" %}

#### Get data from XML

The dataset is being parsed from a stream field xmlUrl that’s being passed on from the ‘Pass URL’ step.
{% endhint %}

1. Drag the ‘Get Data from XML’ step onto the canvas.
2. Create a hop from the ‘Pass URL’ step.
3. Double-click on the step, and configure the following properties:

<figure><img src="/files/LGHUpXn7xI4ImE6nWF1U" alt=""><figcaption><p>Read URL</p></figcaption></figure>

4. Click on the ‘Content’ tab and configure the following properties:

<figure><img src="/files/0iNCbSr2uVEGbBvSLsv9" alt=""><figcaption><p>Select XPath</p></figcaption></figure>

5. Click on the ‘Fields’ tab and configure the following properties:

<figure><img src="/files/8zp1sYLVX2VCUUcmsn45" alt=""><figcaption><p>Configure fields</p></figcaption></figure>

6. Click on the ‘Get Fields’ button.

Next: open the **Dummy** tab.
{% endtab %}

{% tab title="3. Dummy" %}
{% hint style="info" %}

#### Dummy

The Dummy step does not do anything. Its primary function is to be a placeholder for testing purposes. For example, to have a transformation, you need at least two steps connected to each other.
{% endhint %}

1. Drag a ‘Dummy’ step onto the canvas.
2. Create a hop from the ‘Get data from XML’ step.
3. Close the Step.
   {% endtab %}

{% tab title="4. RUN" %}
{% hint style="danger" %}

#### **RUN the Transformation**

Remember to enable the hops and disable the hop in Workflow 1: XML - File

The workflow will fail .. do you know why.?
{% endhint %}

1. Click the Run button in the Canvas Toolbar

<figure><img src="/files/Wp5yh9CSg71fSu0xsT7l" alt="" width="375"><figcaption><p>Invalid data type</p></figcaption></figure>

2. Check the logs.

<figure><img src="/files/D4wwKoQlc9jqIXalyIjt" alt=""><figcaption><p>Logs</p></figcaption></figure>

{% hint style="warning" %}
Looks like Zone data type is alphanumeric (string), not integer.
{% endhint %}

3. Change Zone data type to string and re-run transformation.
4. Click on the Dummy step and Preview data.

<figure><img src="/files/vK9H1JYDED0g3nTNBHJB" alt=""><figcaption><p>Preview Plant Catalog</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# JSON

JavaScript Object Notation ..

{% hint style="info" %}

#### **JSON**

JSON (JavaScript Object Notation) is a lightweight data interchange format that's easy for humans to read and write, and simple for machines to parse and generate. It uses a text-based structure with key-value pairs and arrays to represent data. JSON is language-independent and widely used for transmitting data in web applications.

Now, to extract key-value pairs from this JSON object in Pentaho Data Integration, you would typically use the "JSON Input" step.

```json
{
  "customer": {
    "id": 1001,
    "name": "John Doe",
    "email": "john.doe@example.com",
    "active": true
  }
}
```

In the JSON Input step, the data stream field name, path and data type are defined.
{% endhint %}

| Name   | Path              | Type    |
| ------ | ----------------- | ------- |
| id     | $.customer.id     | Integer |
| name   | $.customer.name   | String  |
| email  | $.customer.email  | email   |
| active | $.customer.active | Boolean |

{% hint style="info" %}
Here's a brief explanation of the JSON Path notation used:

`$` represents the root of the JSON document

`.customer` navigates to the "customer" object

`.id`, `.name`, `.email`, and `.active` access the respective fields within the "customer" object
{% endhint %}

***


# Read JSON

Read JSON objects.

{% hint style="warning" %}

#### Workshop - Read JSON

Parse a JSON file into rows. Use **JSON Input**.

**What you’ll do**

* Read a JSON file from disk.
* Set a loop path for an array.
* Extract fields with JSONPath.
* Use **Get Fields** to infer metadata.
* Preview rows and validate types.

**Prerequisites:** Basic transformations. Basic JSON (objects, arrays). PDI installed.

**Estimated time:** 15 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/992afd2bbf75465ab70ae76d895ea4f3?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
JSON Input
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/KuMmOmVib28mfo1DWknO" %}

***

<figure><img src="/files/x6wBpjxiEX4lzpV0aNaa" alt="" width="375"><figcaption><p>JSON input</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

1. Open `jsonfile.js` in an editor. You will extract fields from this array:

```json
{ "document": {
    "order": [ 
      { "productline": "Classic Cars",
        "customer": "Christine Loomis",
        "status": "Delivered",
        "date": "January 2004",
        "value": 21.99
      },
      { "productline": "Classic Cars",
        "customer": "Mary L. Peachin",
        "status": "Delivered",
        "date": "November 2008",
        "value": 24.99
      },
      { "productline": "Trains",
        "customer": "Bob Italia",
        "status": "Delivered",
        "date": "July 1994",
        "value": 14.99
      }
    ]
  }
}
```

{% hint style="info" %}
`status` can be `Delivered` or `Returned`.
{% endhint %}

***

{% tabs %}
{% tab title="1. JSON Input" %}
{% hint style="info" %}

#### JSON Input

JSON Input reads JSON and outputs rows.

You set one loop path. It outputs one row per loop element.
{% endhint %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag **JSON Input** onto the canvas.
3. Open the step.
4. On the **File** tab, select your `jsonfile.js`.

{% hint style="info" %}
Save your transformation near `jsonfile.js`. Then use a portable path.
{% endhint %}

Example:

```
${Internal.Transformation.Filename.Directory}/jsonfile.js
```

5. Set **Loop path** to:

```
$.document.order[*]
```

6. Open the **Fields** tab.
7. Select **Get Fields**.
8. Verify these field paths and types:

* `productline` (String)
* `customer` (String)
* `status` (String)
* `date` (String)
* `value` (Number)

<figure><img src="/files/XogGI4my2DluGDO2bTpI" alt=""><figcaption><p>JSON input - file</p></figcaption></figure>

<figure><img src="/files/ykbo8iLoe9sIScKRllp2" alt=""><figcaption><p>JSON input - fields</p></figcaption></figure>

{% hint style="info" %}
Need help writing JSONPath? Use a tester like: <https://jsonpath.com/>
{% endhint %}
{% endtab %}

{% tab title="2. Dummy" %}
{% hint style="info" %}

#### Dummy

Dummy does not do anything. Use it as a preview target.
{% endhint %}

1. Expand **Flow** in the Design palette.
2. Drag **Dummy** onto the canvas.
3. Create a hop from **JSON Input** to **Dummy**.
   {% endtab %}

{% tab title="3. Run" %}
{% hint style="info" %}

#### Run the transformation

Run locally and preview the output rows.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. After it finishes, right-click **Dummy**.
3. Select **Preview**.

{% hint style="success" %}
You should see one row per order in the JSON array.
{% endhint %}

<figure><img src="/files/hwTgcdfRdFLe1MJqeXmT" alt=""><figcaption><p>Preview data</p></figcaption></figure>
{% endtab %}
{% endtabs %}

***

### Troubleshooting

<details>

<summary>No rows returned</summary>

Check the **Loop path** first. For this file, it must point to the array:

```
$.document.order[*]
```

If the JSON structure changes, update the loop path.

</details>

<details>

<summary>Fields are null</summary>

Confirm field paths match the JSON keys. If you loop over `order[*]`, use `productline`, not `$.document.order.productline`.

</details>

<details>

<summary>Wrong data types</summary>

Use **Get Fields** as a starting point. Then set types explicitly.

Example: keep `status` as **String**.

</details>


# RSS Feed

RSS reader ..

{% hint style="danger" %}
This step does not work in Pentaho Data Integration 9.5+
{% endhint %}

Steel Wheels have several JSON data sources. In this guided demonstration, you will create a simple workflow to extract the required reporting dataset.

In this guided demonstration, you will configure:

* RSS Input
* Filter Step

RSS (Rich Site Summary; originally RDF Site Summary; often called Really Simple Syndication) uses a family of standard web feed formats to publish frequently updated information: blog entries, news headlines, audio, video.

<figure><img src="/files/hx22wtrkfN949CWRlLlC" alt=""><figcaption><p>RSS</p></figcaption></figure>

#### To create a new transformation

1. In Spoon, click File > New > Transformation:

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking New, then Transformation
* By using the CTRL-N hot key

{% tabs %}
{% tab title="1. RSS Input" %}
{% hint style="info" %}
This step imports data from an RSS or Atom feed. RSS versions 0.91, 0.92, 1.0, 2.0, and Atom versions 0.3 and 1.0 are supported.
{% endhint %}

1. Drag the ‘RSS Input’ step onto the canvas.
2. Double-click on the step, and configure the following properties:
   {% endtab %}

{% tab title="2. Filter rows" %}
{% hint style="info" %}
The Filter Rows step allows you to filter rows based on conditions and comparisons. Once this step is connected to a previous step (one or more and receiving input), you can click on the "", "=" and "" areas to construct a condition.

To enter an IN LIST operator, use a string value separated by semicolons. This also works on numeric values like integers. The list of values must be entered with a string type, e.g.: 2;3;7;8
{% endhint %}

1. Drag the ‘Filter rows’ step onto the canvas.
2. Double-click on the step, and configure the following properties:
   {% endtab %}

{% tab title="3. Text File ouput" %}
{% hint style="info" %}
The Text file output step is used to export data to text file format. This is commonly used to generate Comma Separated Values (CSV files) that can be read by spreadsheet applications. It is also possible to generate fixed width files by setting lengths on the fields in the fields tab.
{% endhint %}

1. Drag the ‘Text File Output’ step onto the canvas.
2. Double-click on the step, and configure the following properties:
   {% endtab %}

{% tab title="4. RUN" %}

1. Click the Run button in the Canvas Toolbar
2. Click on the Text File Output step and Preview data.
   {% endtab %}
   {% endtabs %}


# Databases

Steel Wheels ..

{% hint style="info" %}

#### Steel Wheels - sampledata

Steel Wheels utilizes a straightforward Enterprise Resource Platform (ERP) to manage various Business Units (BUs) including Human Resources, Marketing, Finance, Supply Chain, and others. The upcoming workshops will demonstrate steps that illustrate CRUD (Create, Read, Update, Insert, Delete) operations:
{% endhint %}

<figure><img src="/files/s2PTZjFxgE5OgJNjR6dj" alt=""><figcaption><p>Steel Wheels - sampledata ERP</p></figcaption></figure>

{% hint style="info" %}

#### **Steel Wheels - Schema**

The Steel Wheels ERP database schema, is designed for a manufacturing or retail company that manages complex product sales and distribution operations. The database centers around order processing, tracking customer purchases through detailed order and order detail records, while maintaining comprehensive customer profiles including contact information, territories, and credit limits.

The system supports inventory management through product tracking with quantities, pricing, and vendor relationships, and includes employee management with sales performance monitoring across different office locations.

Additional analytical capabilities are built in through summary tables for customer orders, monthly sales trends, and departmental reporting, along with specialized views for tracking product performance, payment histories, and trial balances, suggesting this ERP system serves a business that requires detailed financial reporting and sales analytics across multiple regions and product lines.

**Tables**

**OFFICES**: Stores company office locations with address details

**EMPLOYEES**: Contains employee information with relationships to offices and reporting structure

**CUSTOMERS**: Stores customer information including contact details and credit limits

**PRODUCTS**: Contains product catalog with inventory and pricing information

**ORDERS**: Tracks customer orders with status and dates

**ORDERDETAILS**: Contains line items for each order with quantity and price

**PAYMENTS**: Records customer payments with amounts and dates

**ORDERFACT**: A fact table for order analytics

**CUSTOMER\_W\_TER**: Extended customer information with territory

**DIM\_TIME**: Time dimension table for reporting

**DEPARTMENT\_MANAGERS**: Stores department manager information

**QUADRANT\_ACTUALS**: Contains budget vs. actual financial data with a generated VARIANCE column

**TRIAL\_BALANCE**: Financial accounting data

**Views**

**customer\_order\_summary**: Summarizes orders and spending by customer

**product\_performance**: Analyzes product sales metrics including revenue and profit

**employee\_sales\_performance**: Tracks sales performance by employee

**monthly\_sales\_trend**: Shows sales trends over time by month

**product\_inventory\_status**: Categorizes products by inventory levels

**customer\_payment\_history**: Summarizes customer payment activity and balances

**Stored Procedures**

**GetCustomerOrders**: Retrieves orders for a specific customer

**UpdateProductStock**: Updates product inventory levels

**GetProductSalesByQuarter**: Analyzes quarterly product sales

**GetTopCustomersByRegion**: Identifies top customers by region

**GetInventoryValueByProductLine**: Calculates inventory metrics by product line

**Triggers**

**before\_order\_insert**: Validates date constraints on orders

**before\_payment\_insert**: Ensures payment amounts are positive

**Events**

* **daily\_maintenance**: Scheduled task for database maintenance
  {% endhint %}

***


# CRUID

CRUID database operations are a set of five basic functions that allow us to manipulate data in a persistent storage system, such as a relational database ..

{% hint style="info" %}

#### Database Operations

Steel Wheels utilizes a straightforward Enterprise Resource Platform (ERP) to manage various Business Units (BUs) including Human Resources, Marketing, Finance, Supply Chain, and others. The upcoming workshops will demonstrate steps that illustrate CRUD (Create, Read, Update, Insert, Delete) operations:

**Create** operations add new records to a database. This might involve inserting a new customer profile, product listing, or transaction record. In SQL, this is typically done using the INSERT statement, while in NoSQL databases, it might use methods like insertOne() or save().

**Read** operations retrieve existing data from the database. This could be fetching a single record by its unique identifier or querying multiple records based on specific criteria. SQL uses SELECT statements for this purpose, while NoSQL databases might use find() or get() methods.

**Update** operations modify existing records in the database. This could involve changing a customer's address, updating a product's price, or modifying any stored information. SQL uses the UPDATE statement, while NoSQL databases might use methods like updateOne() or save() on an existing document.

**Delete** operations remove records from the database. This might involve permanently removing a user account or archiving old data. SQL uses the DELETE statement, while NoSQL databases typically use methods like deleteOne() or remove().
{% endhint %}

<figure><img src="/files/M7DbLvd4At1YGa4QwnnR" alt=""><figcaption><p>CRUID</p></figcaption></figure>


# Database Connections

Database connections ..

{% hint style="warning" %}

#### Workshop - Database connections

Create a reusable MySQL connection to the Steel Wheels `sampledata` database.

**What you’ll do**

* Validate the database is reachable (optional, using DBeaver)
* Install a JDBC driver (only if PDI does not include it)
* Create, test, share, and explore a PDI database connection

**Prerequisites**

* PDI (Spoon) installed and working
* A running `sampledata` database (Docker setup recommended)
* Basic understanding of schemas, tables, and authentication

**Estimated time:** 15 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/3ec2d5123814460d92085851c18daaee?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Database Connection
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download the following file.

Keep the filename unchanged.

Save it in your workshop folder.
{% endhint %}

{% file src="/files/3ILjuwjLO9JaUpUsPHpq" %}

***

{% hint style="info" %}

### Create a new transformation

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. DBeaver" %}
{% hint style="info" %}

#### **DBeaver**

DBeaver is optional.\
Use it to confirm the database is reachable before you touch PDI.

DBeaver CE ships with most drivers.\
That makes it a fast connectivity check.
{% endhint %}

{% hint style="warning" %}
Pentaho no longer ships a writable sample database user by default.\
Use the Docker `sampledata` setup for hands-on database workshops.
{% endhint %}

{% tabs %}
{% tab title="MySQL" %}
{% hint style="info" %}

#### **MySQL Database**

If you completed the [Setup](broken://spaces/ZpCSy6Skj215f4oWypdc/pages/nyuNuY53XaIkQsM4MkZX), you should have a MySQL Docker container.\
It should be exposed on port `3306` and include the `sampledata` database.
{% endhint %}

1. Launch DBeaver and select **MySQL**.

<figure><img src="/files/f04HxAgnFLcJJfT7g1Kc" alt=""><figcaption><p>MySQL</p></figcaption></figure>

2. Configure the connection:

* **Username:** `root`
* **Password:** `password`

<figure><img src="/files/YlrpU0j9USa8MQwptL9F" alt=""><figcaption><p>Configure &#x26; Test MySQL connection - sampledata</p></figcaption></figure>

{% hint style="warning" %}
You might need to download the supported driver version.

If the test fails, enable `allowPublicKeyRetrieval`.
{% endhint %}

<figure><img src="/files/8NYNZW9GthDyeTa13Vdc" alt=""><figcaption><p>Enable: allowPublicKeyRetrieval</p></figcaption></figure>

3. Test the connection.

<figure><img src="/files/zi3B824EcsskmFjDo7Am" alt=""><figcaption><p>Test connection</p></figcaption></figure>

4. Expand **Databases** > **sampledata** > **Tables**.

<figure><img src="/files/IpfBGicoVihojG7VvsCo" alt=""><figcaption><p>Customer Data</p></figcaption></figure>

5. Open a SQL window and run a test query:

```sql
select * from CUSTOMERS
where COUNTRY = 'USA' and CITY = 'NYC';
```

<figure><img src="/files/5B25msWIcQHRygRLGOTv" alt=""><figcaption><p>Sql query - NYC Customers</p></figcaption></figure>

{% hint style="success" %}
Checkpoint: you can browse tables and run a query against `CUSTOMERS`.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. JDBC Driver" %}
{% hint style="info" %}

#### **Download JDBC Driver**

PDI does not ship all JDBC drivers.\
If your database type is missing, add the driver JAR.
{% endhint %}

{% embed url="<https://docs.pentaho.com/install/jdbc-drivers-reference>" %}

1. Download the JDBC driver for your database.

{% embed url="<https://dbschema.com/databases.html>" %}

2. Copy the driver JAR into your PDI install:

{% tabs %}
{% tab title="Windows" %}
`C:\Pentaho\design-tools\data-integration\lib\`
{% endtab %}

{% tab title="macOS / Linux" %}
`~/Pentaho/design-tools/data-integration/lib/`
{% endtab %}
{% endtabs %}

3. Restart Spoon.

{% hint style="info" %}
If your install uses `lib/jdbc/`, place the JAR there instead.
{% endhint %}
{% endtab %}

{% tab title="3. Data Integration" %}
{% hint style="info" %}

#### **Pentaho Data Integration Connection**

Create the connection once.\
Reuse it in steps like **Table input**, **Table output**, and **Database lookup**.

In this lab, you connect to the Steel Wheels `sampledata` database (MySQL).
{% endhint %}

**Define a database connection (MySQL)**

1. Create a transformation.
2. In Spoon, select **File** > **New** > **Database connection**.

The **Database connection** dialog opens.

<figure><img src="/files/nfLYAjgnNcRbRwjO96lX" alt=""><figcaption></figcaption></figure>

3. Enter the following details:

{% hint style="danger" %}
If you use a MariaDB driver newer than `2.7.x`, you might see a **fetch size** error.\
If that happens, use the **MySQL** driver instead.
{% endhint %}

* **Connection name:** `MySQL: sampledata`
* **Connection type:** **MySQL**
* **Access:** **Native (JDBC)**
* **Host name:** `localhost` (or your Docker host IP)
* **Database name:** `sampledata`
* **Username:** `pentaho_admin`
* **Password:** `password`

<figure><img src="/files/QyouGLeHqt30IXT0SXd8" alt=""><figcaption><p>MySQL - sampledata</p></figcaption></figure>

4. Select **Test**.

{% hint style="info" %}
Checkpoint: Spoon shows a success message.
{% endhint %}

{% tabs %}
{% tab title="1. Share Connection" %}
{% hint style="info" %}
**Share Database Connection**

Share the connection so other transformations can reuse it.
{% endhint %}

1. Click OK to save your entries and exit the Database Connection dialog box.
2. From within the View tab, right-click on the connection and select Share from the list that appears.

<figure><img src="/files/McZz5Iv5N64RmHHYT76Z" alt=""><figcaption><p>Share database connection</p></figcaption></figure>

{% hint style="info" %}
Shared connections show up for other users and projects.\
Use **Explore** to confirm schemas and tables.
{% endhint %}
{% endtab %}

{% tab title="2. Explore Database" %}
{% hint style="info" %}
**Explore Database**

Use **Database Explorer** to browse schemas, preview rows, and run SQL.
{% endhint %}

1. Click on the View tab, expand Database Connections.
2. Right-click MySQL:sampledata and choose Explore from the menu options:

|                                   | Action                                                                |
| --------------------------------- | --------------------------------------------------------------------- |
| Preview the first 100 rows of ..  | Return the first 100 rows of the selected table.                      |
| Preview first .. rows of ..       | Enter the number of rows to preview                                   |
| Number of rows ..                 | Displays number of rows                                               |
| Generate DDL                      | Displays DDL statement that creates table.                            |
| Generate DDL for other connection | Select connection to display DDL. Syntax is based on database engine. |
| Open SQL for ..                   | Edit SELECT statement                                                 |
| Truncate table                    | Deletees all the rows from selected table                             |

3. In the Database Explorer window, expand Sampledata > Tables

<figure><img src="/files/oOLOxwr8SdpjHhn62BG8" alt="" width="375"><figcaption><p>Database Explorer - sampledata</p></figcaption></figure>

4. Right-click the `CUSTOMERS` table and choose **Preview first 100**.
5. Examine the customer data.
6. Select **View SQL**.

<figure><img src="/files/1z3WddQJPibQc3251vHf" alt="" width="375"><figcaption><p>SQL</p></figcaption></figure>

7. Click Execute.

<figure><img src="/files/6qB4vJCoFdwQd8x7OxBi" alt=""><figcaption><p>Execute SQL statement</p></figcaption></figure>

{% hint style="success" %}
Checkpoint: you can preview `CUSTOMERS` and execute SQL in Database Explorer.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Create DB table

Create tables ..

{% hint style="warning" %}

#### Workshop - Create DB table

Create a table from a stream definition.\
Load rows into that table in the same run.

**What you’ll do**

* Read a delimited file into a stream
* Map stream fields to table columns
* Generate and run `CREATE TABLE` SQL from **Table Output**
* Insert rows with commit and batch settings

**Prerequisites**

* A working database connection. See [Database Connections](/pentaho-data-integration/data-integration/data-sources/databases/cruid/database-connections).
* Basic understanding of tables and SQL data types

**Estimated time:** 30 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/ebcc69cd2a9347f8bec1620259952df7?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/GABq2FDI1hWBtd8TM1rK" %}

{% file src="/files/laqnH0bKB6FdtKFs3wi3" %}

{% file src="/files/VMjJDhgZTsjGuFmPHSiN" %}

***

<figure><img src="/files/1IJE894zsoO3fgK4yV0P" alt="" width="375"><figcaption><p>Create databases</p></figcaption></figure>

{% hint style="info" %}

### Create a new transformation

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="Workflow 1 - Sales Data" %}
{% hint style="info" %}

#### **Load Sales Data**

Load sales data from a CSV file into `STG_SALES_DATA`.

Adjust field lengths to avoid truncation.
{% endhint %}

<figure><img src="/files/Erbl7dFoZFgdZThpPMCo" alt="" width="277"><figcaption><p>Load sales data</p></figcaption></figure>

Follow the steps outlined below:

{% tabs %}
{% tab title="1. CSV File Input" %}
{% hint style="info" %}
**CSV File Input**

Read a delimited file into a stream.\
Use **Text File Input** if you need more format options.
{% endhint %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
**Start Spoon**

{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag the CSV file input step onto the canvas.
3. Open the CSV file input properties dialog box.

Configure these key fields:

* **Step name:** `csvi-sales_data`
* **File name:** `${Internal.Transformation.Filename.Directory}/sales_data.csv`
* **Delimiter:** `,` (comma)
* **Lazy conversion:** clear
* **Header row present:** select

{% hint style="info" %}
If `${Internal.Transformation.Filename.Directory}` is empty, save the transformation first.
{% endhint %}

{% hint style="danger" %}
CSV File Input infers field lengths from a sample.\
Increase string lengths before you generate table DDL.
{% endhint %}

4. Ensure the following details are configured, as outlined below:

<figure><img src="/files/zOGpA8O7STKx8xze795J" alt=""><figcaption><p>CSV file input</p></figcaption></figure>

5. Click on the Get Fields button.
6. Select **OK**.
   {% endtab %}

{% tab title="2. Table Output" %}
{% hint style="info" %}
**Table Output**

Load rows into a database table.\
This step uses SQL `INSERT`.
{% endhint %}

1. Drag the Table Output step onto the canvas.
2. Open the Table Output properties dialog box. Ensure the following details are configured, as outlined below:

<figure><img src="/files/i37VQPS6Abk4BNCSvSsX" alt=""><figcaption><p>Table output - options</p></figcaption></figure>

3. Click on the Database fields.
4. Click on the ‘Get Fields’ button.

<figure><img src="/files/4hAWwn0z1XNExKfCa6z9" alt=""><figcaption><p>Table output - fields</p></figcaption></figure>

{% hint style="warning" %}
Confirm the mappings between **Table fields** and **Stream fields**.
{% endhint %}

5. Click on the SQL button.

<figure><img src="/files/qgYtFQ35xSmpVJix96pl" alt="" width="375"><figcaption><p>SQL editor</p></figcaption></figure>

6. Select **Execute**.
7. Select **OK** to close all dialogs.

{% hint style="success" %}
Checkpoint: `STG_SALES_DATA` exists in the database.
{% endhint %}
{% endtab %}

{% tab title="3. RUN" %}
{% hint style="info" %}
**RUN and validate**

Use a database tool to verify results (DBeaver, Workbench, or your IDE).\
In production, you usually manage DDL with migrations or scripts.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Confirm the table exists in your `sampledata` database.

<figure><img src="/files/w2IODoGL4Jxn1V4OAUyM" alt=""><figcaption><p>STG_SALES_DATA</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Workflow 2 - Orders" %}
{% hint style="info" %}

#### **Load Orders**

Load orders data from a delimited file into `STG_ORDERS_MERGED`.

Adjust field lengths to avoid truncation.
{% endhint %}

<figure><img src="/files/fTixivdeazDfoVP0lire" alt="" width="288"><figcaption><p>Load orders data</p></figcaption></figure>

Follow the steps outlined below:

{% tabs %}
{% tab title="1. CSV File Input" %}
{% hint style="info" %}
**CSV File Input**

Same setup as Workflow 1, but point to your orders file.
{% endhint %}

1. Drag the CSV file input step onto the canvas.
2. Open the CSV file input properties dialog box.

Ensure the following details are configured, as outlined below:

<figure><img src="/files/9STa2CDOhvSF9khcd5OC" alt=""><figcaption><p>Add path to orders.txt</p></figcaption></figure>

<figure><img src="/files/j692ELltSo63GQWNFPxC" alt=""><figcaption><p>Set Content</p></figcaption></figure>

<figure><img src="/files/tXyHrT4vy7E9Xui34ype" alt=""><figcaption><p>Get Fields</p></figcaption></figure>
{% endtab %}

{% tab title="2. Table Output" %}
{% hint style="info" %}
**Table Output**

Load rows into a database table.\
This step uses SQL `INSERT`.
{% endhint %}

1. Drag the Table Output step onto the canvas.
2. Open the Table Output properties dialog box. Ensure the following details are configured, as outlined below:

<figure><img src="/files/z0juo75eYE6QoIZQAthN" alt=""><figcaption><p>Table output - options</p></figcaption></figure>

3. Click on the Database fields.
4. Click on the ‘Get Fields’ button.

<figure><img src="/files/jVc7QAh1QsVUOpJDwFgY" alt=""><figcaption><p>Table output -fields</p></figcaption></figure>

5. Click on the SQL button.

<figure><img src="/files/KuNskbRAnAupIALHLfFK" alt=""><figcaption><p>SQL editor</p></figcaption></figure>

6. Click Execute.
7. Select **OK** to close all dialogs.

{% hint style="success" %}
Checkpoint: `STG_ORDERS_MERGED` exists in the database.
{% endhint %}
{% endtab %}

{% tab title="3. RUN" %}
{% hint style="info" %}
**RUN and validate**

Use a database tool to verify results.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Confirm the table exists in your `sampledata` database.

<figure><img src="/files/sp2ghOBFEUzhGHIw3lsa" alt=""><figcaption><p>STG_ORDERS_MERGED</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}

<details>

<summary>Troubleshooting</summary>

**“Table already exists”**\
Disable table creation in **Table Output**, or drop the table first.

**String truncation / data too long**\
Increase field lengths in **CSV File Input** before you generate SQL.

**SQL window is empty**\
Select **Get fields** on the **Database fields** tab first.

</details>


# Read DB table

Read shipped orders from a database table using Table Input.

{% hint style="warning" %}

#### Workshop - Read DB table

Build a transformation that reads `ORDERS` rows from a database.\
Filter to shipped orders, calculate lead time, and label late shipments.

**What you’ll do**

* Read rows with **Table Input**
* Generate SQL with **Get SQL select statement**
* Add a calculated field with **Calculator**
* Bucket values with **Number range**
* Sort and format output for review

**Prerequisites**

* Pentaho Data Integration installed and configured
* A working database connection. See [Database Connections](/pentaho-data-integration/data-integration/data-sources/databases/cruid/database-connections).
* Basic `SELECT` and `WHERE`

**Estimated time:** 20 minutes
{% endhint %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/EyMpG5AvaxL66l4wJtEY" %}

***

<figure><img src="/files/k7s9bmUTrUTiyykuH8Yc" alt="" width="563"><figcaption><p>Read from a database</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Table input" %}
{% hint style="info" %}

#### Table Input

Read rows from a database using a connection and SQL.\
In this workshop, filter to orders with `STATUS = 'Shipped'`.
{% endhint %}

1. Start Spoon.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag **Table Input** onto the canvas.
3. Open the step properties.
4. Configure the step to match your environment:
   * Select your database **Connection**
   * Use **Get SQL select statement** to generate a base query
   * Add a `WHERE` clause for shipped orders

{% hint style="info" %}
Example filter:

```sql
WHERE STATUS = 'Shipped'
```

{% endhint %}

<figure><img src="/files/Mp95ZeWETayxAqb1pHjL" alt=""><figcaption><p>Table input</p></figcaption></figure>

5. Select **Preview**. Confirm you get shipped orders.
6. Select **OK**.

{% hint style="success" %}
Checkpoint: Preview shows only rows with `STATUS = 'Shipped'`.
{% endhint %}
{% endtab %}

{% tab title="2. Calculator" %}
{% hint style="info" %}

#### Calculator

Add a derived field using built-in functions.\
Use Calculator for speed and simple expressions.
{% endhint %}

1. Add a hop from **Table Input** to **Calculator**.
2. Drag **Calculator** onto the canvas.
3. Open the step properties.
4. Configure the calculation shown in the screenshot.

<figure><img src="/files/qashpiRTvBF5fhrldrKg" alt=""><figcaption><p>Calculate diff days</p></figcaption></figure>

5. Select **OK**.

{% hint style="info" %}
This creates `order_time`.\
It represents the day difference between required and shipped dates.
{% endhint %}
{% endtab %}

{% tab title="3. Number range" %}
{% hint style="info" %}

#### Number range

Map numeric values into named buckets.\
This makes reports easier to scan.
{% endhint %}

1. Add a hop from **Calculator** to **Number range**.
2. Drag **Number range** onto the canvas.
3. Open the step properties.
4. Configure the ranges as shown.

<figure><img src="/files/wdFugYWCXqNYkNY7nLDz" alt=""><figcaption><p>Number range</p></figcaption></figure>

{% hint style="info" %}
This writes an output label (for example, `order_status`) based on `order_time`.\
Use the same labels and thresholds as the screenshot.
{% endhint %}

5. Select **OK**.
   {% endtab %}

{% tab title="4. Sort rows" %}
{% hint style="info" %}

#### Sort rows

Sort output to match how you want to read it.\
This is also a common prerequisite for merge-style steps.
{% endhint %}

1. Add a hop from **Number range** to **Sort rows**.
2. Drag **Sort rows** onto the canvas.
3. Open the step properties.
4. Configure the sort keys as shown.

<figure><img src="/files/9um1SHb23k5k5dEy1uiV" alt=""><figcaption><p>Sort rows</p></figcaption></figure>

5. Select **OK**.

{% hint style="info" %}
If you hit memory errors, lower the sort size.\
PDI spills to temp files when needed.
{% endhint %}
{% endtab %}

{% tab title="5. Select values" %}
{% hint style="info" %}

#### Select values

Keep only fields you need.\
Fix types, lengths, and formats for downstream steps.
{% endhint %}

1. Add a hop from **Sort rows** to **Select values**.
2. Drag **Select values** onto the canvas.
3. Open the step properties.
4. Configure the field selection and type changes shown.
5. Select **OK**.

{% hint style="info" %}
This formats `REQUIREDDATE` and `SHIPPEDDATE`.
{% endhint %}
{% endtab %}

{% tab title="6. RUN" %}
{% hint style="info" %}

#### Run and validate

Run the transformation and inspect the final stream.
{% endhint %}

1. In Spoon, select **Run**.
2. In **Execution Results**, open **Preview data** for **Select values**.

<figure><img src="/files/35FTE8eV7PdYe1CUbXAL" alt=""><figcaption><p>Status of 'shipped' orders</p></figcaption></figure>

{% hint style="success" %}
Checkpoint: You see shipped orders plus your derived fields (`order_time`, and the range label).
{% endhint %}
{% endtab %}
{% endtabs %}

<details>

<summary>Troubleshooting</summary>

**Preview shows zero rows**\
Confirm the `WHERE STATUS = 'Shipped'` filter matches your source values.

**SQL errors**\
Select **Get SQL select statement** again. Then re-apply your `WHERE` clause.

**Date or number conversion issues**\
Fix types in **Select values**. Re-run the preview.

**Out of memory during sort**\
Lower the sort size in **Sort rows**, or increase JVM memory.

</details>


# Update DB table

Update employees in the EMPLOYEES table using the Update step.

{% hint style="warning" %}

#### Workshop - Update DB table

Update existing rows in `EMPLOYEES` using a key lookup.\
You will change job titles for two employees.

**What you’ll do**

* Read an update file with **Text file input**
* Update matching rows with **Update**
* Validate changes with SQL

**Prerequisites**

* A working database connection. See [Database Connections](/pentaho-data-integration/data-integration/data-sources/databases/cruid/database-connections).
* Basic primary key concepts (`EMPLOYEENUMBER`)

**Estimated time:** 20 minutes
{% endhint %}

***

{% hint style="info" %}
**Workshop files**

Create a file named `employees_update.txt`.\
Save it in the same folder as your transformation.

This workshop assumes a comma-delimited file with **no header row**:

```
1002,Murphy,Diane,x5800,dmurphy@classicmodelcars.com,1,1000,CEO
1102,Bondur,Gerard,x5408,athompson@classicmodelcars.com,4,1056,Regional Sales Manager (EMEA)
```

{% endhint %}

***

<figure><img src="/files/luKDRk7sJP1sC6OUE0jF" alt="" width="563"><figcaption><p>Update employees</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Text file input" %}
{% hint style="info" %}

#### Text file input

Read the incoming updates from `employees_update.txt`.
{% endhint %}

1. Start Spoon.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag **Text file input** onto the canvas.
3. Open the step properties.
4. Configure the file path:
   * **File:** `${Internal.Transformation.Filename.Directory}/employees_update.txt`

<figure><img src="/files/XAJFHvePsYw5iO6oZRqf" alt=""><figcaption><p>Set file path</p></figcaption></figure>

{% hint style="info" %}
If `${Internal.Transformation.Filename.Directory}` is empty, save the transformation first.
{% endhint %}

5. Select **Content**. Use the same delimiter settings as the screenshot.

<figure><img src="/files/yAxiIQhIuZwGTBxIY9yV" alt=""><figcaption><p>Text file input - Content</p></figcaption></figure>

6. Select **Get Fields**.
7. On **Fields**, confirm you have these stream fields:
   * `EMPLOYEE_NUMBER`
   * `LASTNAME`
   * `FIRSTNAME`
   * `EXTENSION`
   * `EMAIL`
   * `OFFICECODE`
   * `REPORTSTO`
   * `JOBTITLE`

<figure><img src="/files/Z3nai2uadnbxLoYj0Tni" alt=""><figcaption><p>Text file input - Fields</p></figcaption></figure>

8. Optional: select **Preview**. Confirm you get 2 rows.
9. Select **OK**.

{% hint style="success" %}
Checkpoint: Preview shows 2 employee rows.
{% endhint %}
{% endtab %}

{% tab title="2. Update" %}
{% hint style="info" %}

#### Update

Use **Update** to update existing database rows only.\
If a key lookup does not match, the step skips that row.
{% endhint %}

{% hint style="info" %}
If you also need inserts, use [Insert / Update DB](/pentaho-data-integration/data-integration/data-sources/databases/cruid/insert-update-db).
{% endhint %}

1. Drag **Update** onto the canvas.
2. Create a hop from **Text file input** to **Update**.
3. Open the step properties.
4. Select your database **Connection**.
5. Set **Target table** to `EMPLOYEES`.

<figure><img src="/files/LvzhLxcweAoUdlO1ofb8" alt="" width="563"><figcaption><p>Update fields</p></figcaption></figure>

{% hint style="info" %}
**Key lookup**

Map the table key `EMPLOYEENUMBER` to the stream field `EMPLOYEE_NUMBER`.

**Update fields**

Select **Get update fields**. Then confirm mappings are correct.

Do not add `EMPLOYEENUMBER` as an update field.
{% endhint %}

6. Select **OK**.
   {% endtab %}

{% tab title="3. Run and validate" %}
{% hint style="info" %}

#### Run and validate

Run the transformation. Then validate the changes in the database.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. In **Execution Results**, open **Step Metrics**.

<figure><img src="/files/IFb6orkG7JKxnaAC0SoA" alt=""><figcaption><p>Step metrics</p></figcaption></figure>

{% hint style="info" %}
You should see **2 updated** rows for the Update step.
{% endhint %}

3. Verify the updated rows:

```sql
select * from EMPLOYEES
where EMPLOYEENUMBER in ('1002','1102');
```

<figure><img src="/files/GJRaZMM1lw5q5mhFX8Wa" alt=""><figcaption><p>Update employees</p></figcaption></figure>

{% hint style="success" %}
Checkpoint: `JOBTITLE` matches the values from `employees_update.txt`.
{% endhint %}
{% endtab %}
{% endtabs %}

<details>

<summary>Troubleshooting</summary>

**Step updates 0 rows**\
Confirm `EMPLOYEENUMBER` exists in `EMPLOYEES`. The Update step does not insert.

**Updates fail with data type errors**\
Make sure `EMPLOYEE_NUMBER` is numeric. Use **Select values** to cast if needed.

**Wrong rows updated**\
Confirm the key mapping is `EMPLOYEENUMBER` (table) = `EMPLOYEE_NUMBER` (stream).

</details>


# Insert / Update DB

Insert new employees and update existing employees using the Insert/Update step.

{% hint style="warning" %}

#### Workshop - Insert / Update DB

Synchronize a file feed to the `EMPLOYEES` table.\
Some rows already exist. Others are new.

This pattern is called **upsert**: update if found, insert if not.

**What you’ll do**

* Read a mixed employee feed with **Text file input**
* Upsert into `EMPLOYEES` with **Insert/Update**
* Validate inserted and updated rows with SQL

**Prerequisites**

* A working database connection. See [Database Connections](/pentaho-data-integration/data-integration/data-sources/databases/cruid/database-connections).
* Basic primary key concepts (`EMPLOYEENUMBER`)

**Estimated time:** 20 minutes
{% endhint %}

***

{% hint style="info" %}
**Workshop files**

Create a file named `employees_insert_update.txt`.\
Save it in the same folder as your transformation.

This workshop assumes a comma-delimited file with **no header row**:

```
1188,Firrelli,Julianne,x2174,jfirrelli@classicmodelcars.com,2,1143,Sales Manager
1619,King,Tom,x6324,tking@classicmodelcars.com,6,1088,Sales Rep
1810,Lundberg,Anna,x910,alundberg@classicmodelcars.com,2,1143,Sales Rep
1811,Schulz,Chris,x951,cschulz@classicmodelcars.com,2,1143,Sales Rep
```

{% endhint %}

***

<figure><img src="/files/16uwUlDrmqMvKQ1DdZkL" alt="" width="375"><figcaption><p>Insert / Update</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Text file input" %}
{% hint style="info" %}

#### Text file input

Read the incoming mixed feed from `employees_insert_update.txt`.
{% endhint %}

1. Start Spoon.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag **Text file input** onto the canvas.
3. Open the step properties.
4. Configure the file path:
   * **File:** `${Internal.Transformation.Filename.Directory}/employees_insert_update.txt`

<figure><img src="/files/SQEtBgMli03zzbP0Kx4D" alt=""><figcaption><p>Text File input - File</p></figcaption></figure>

{% hint style="info" %}
If `${Internal.Transformation.Filename.Directory}` is empty, save the transformation first.
{% endhint %}

5. Select **Content**. Use the same delimiter settings as the screenshot.

<figure><img src="/files/yAxiIQhIuZwGTBxIY9yV" alt="" width="563"><figcaption><p>Text file input - Content</p></figcaption></figure>

6. Select **Get Fields**.
7. On **Fields**, confirm you have these stream fields:
   * `EMPLOYEE_NUMBER`
   * `LASTNAME`
   * `FIRSTNAME`
   * `EXTENSION`
   * `EMAIL`
   * `OFFICECODE`
   * `REPORTSTO`
   * `JOBTITLE`

<figure><img src="/files/TciqAYNerv1GcdfZQrqH" alt=""><figcaption><p>Text File input - Fields</p></figcaption></figure>

8. Optional: select **Preview**. Confirm you get 4 rows.
9. Select **OK**.

{% hint style="success" %}
Checkpoint: Preview shows 4 employee rows.
{% endhint %}
{% endtab %}

{% tab title="2. Insert / Update" %}
{% hint style="info" %}

#### Insert / Update

Upsert into the `EMPLOYEES` table.
{% endhint %}

1. Drag **Insert / Update** onto the canvas.
2. Create a hop from **Text file input** to **Insert / Update**.
3. Open the step properties.
4. Select your database **Connection**.
5. Set **Target table** to `EMPLOYEES`.

<figure><img src="/files/1fH7KqEfvUXKC0ZjxPwd" alt="" width="563"><figcaption><p>Insert / update options</p></figcaption></figure>

{% hint style="info" %}
**Key lookup**

Map the table key `EMPLOYEENUMBER` to the stream field `EMPLOYEE_NUMBER`.

**Update fields**

Select **Get update fields**. Then confirm mappings are correct.

Do not add `EMPLOYEENUMBER` as an update field.
{% endhint %}

{% hint style="warning" %}
Do not enable **Update the keys**.
{% endhint %}

6. Select **OK**.
   {% endtab %}

{% tab title="3. Run and validate" %}
{% hint style="info" %}

#### Run and validate

Run the transformation. Then validate the results in the database.
{% endhint %}

1. Select **Run** in the canvas toolbar.
2. In **Execution Results**, open **Step Metrics**.

<figure><img src="/files/C8rnknkkOFmEsEspuO6Q" alt=""><figcaption><p>Step metrics</p></figcaption></figure>

{% hint style="info" %}
Expect a mix of inserts and updates, depending on your starting data.\
You should see activity in the Insert/Update step metrics.
{% endhint %}

3. Verify the four employees exist:

```sql
select * from EMPLOYEES
where EMPLOYEENUMBER in ('1188','1619','1810','1811');
```

<figure><img src="/files/93oTUHLwZi9mJRcYD1Ps" alt=""><figcaption><p>Insert / Update Employees</p></figcaption></figure>

{% hint style="success" %}
Checkpoint: All four `EMPLOYEENUMBER` values exist in `EMPLOYEES`.
{% endhint %}
{% endtab %}
{% endtabs %}

<details>

<summary>Troubleshooting</summary>

**All rows updated (no inserts)**\
Those employee numbers already exist in `EMPLOYEES`.

**All rows inserted (no updates)**\
Your `EMPLOYEENUMBER` values were not found, or key mapping is wrong.

**Duplicate key errors**\
You likely used **Table Output** instead of **Insert/Update**, or you mapped keys incorrectly.

**Updates affect too many rows**\
Ensure `EMPLOYEENUMBER` is unique in your table. Fix duplicates before you upsert.

</details>


# Delete DB table

Delete rows from STG\_SALES\_DATA using transformation-driven criteria.

{% hint style="warning" %}

#### Workshop - Delete DB table

Delete rows from `STG_SALES_DATA` based on criteria in a stream.\
This workshop uses a product line list and a minimum quantity threshold.

Use this step when your delete logic is driven by transformation output.\
Use **Execute SQL script** for simple deletes.

**What you’ll do**

* Inspect the target table before you delete
* Build a delete criteria stream (product line + quantity)
* Delete matching rows with **Delete**
* Validate results with SQL

**Prerequisites**

* `STG_SALES_DATA` exists. Create it in [Create DB table](/pentaho-data-integration/data-integration/data-sources/databases/cruid/create-db-table).
* A working database connection. See [Database Connections](/pentaho-data-integration/data-integration/data-sources/databases/cruid/database-connections).

**Estimated time:** 20 minutes
{% endhint %}

***

{% hint style="info" %}
**Workshop files**

Create a file named `productlines.csv`.\
Save it in the same folder as your transformation.

This workshop assumes a single-column file with **no header row**:

```
Classic Cars
Motorcycles
Planes
Ships
Trains
Trucks and Buses
Vintage Cars
```

{% endhint %}

{% hint style="danger" %}
Back up your table before you run deletes.\
You can copy the table, or use a database snapshot.
{% endhint %}

{% file src="/files/GEjwJhmXShQ5dizQjxDb" %}

{% file src="/files/RssdwRBksrQE9M1YUkRz" %}

***

<figure><img src="/files/0C1ipQ0xKZe8bL915cxH" alt="" width="563"><figcaption><p>Delete</p></figcaption></figure>

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

{% tabs %}
{% tab title="1. Inspect data" %}
{% hint style="info" %}

#### Inspect the data

Inspect `STG_SALES_DATA` before you delete.\
You need a baseline to validate the change.
{% endhint %}

1. In your database tool, view `STG_SALES_DATA`.

<figure><img src="/files/A0HZhONjzK0uI0FHZlKd" alt=""><figcaption><p>STG_SALES_DATA</p></figcaption></figure>

{% hint style="info" %}
You will filter deletes using `PRODUCTLINE` and `QUANTITYORDERED`.
{% endhint %}

2. Run a quick check for high-quantity rows:

```sql
select * from STG_SALES_DATA
where QUANTITYORDERED > 50;
```

<figure><img src="/files/LX4dMDBgmknUR9Pc40wU" alt=""><figcaption><p>STG_SALES_DATA constraint QUANTITYORDERED > 50</p></figcaption></figure>
{% endtab %}

{% tab title="2. CSV File input" %}
{% hint style="info" %}

#### CSV File input

Read `productlines.csv`. Each row is one `PRODUCTLINE` value.
{% endhint %}

1. Start Spoon.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

2. Drag the CSV File Input step onto the canvas.
3. Open the CSV File Input properties dialog box.
4. Configure it to read: `${Internal.Transformation.Filename.Directory}/productlines.csv`
5. Select **Get Fields**.

<figure><img src="/files/HLatfd3JlVS8n2inVvdx" alt=""><figcaption><p>CSV File input - PRODUCTLINE list</p></figcaption></figure>
{% endtab %}

{% tab title="3. Parameters" %}
{% hint style="info" %}

#### Transformation parameters

Use a parameter for your minimum quantity threshold.\
This keeps your transformation easy to reuse.
{% endhint %}

1. Double-click on the canvas and select the Parameter tab.
2. Create a parameter named `min_quantityordered`.
3. Set a default value (for example `50`).

<figure><img src="/files/CwMIfLFQYR32mNS7bvfG" alt="" width="563"><figcaption><p>Set parameters</p></figcaption></figure>
{% endtab %}

{% tab title="4. Get Variables" %}
{% hint style="info" %}

#### Get variables

Bring `min_quantityordered` into the stream so the Delete step can use it.
{% endhint %}

1. Drag the Get variables step onto the canvas.
2. Open the Get variables properties dialog box.
3. Configure the step to output a field for `${min_quantityordered}`.

<figure><img src="/files/Wt6ffHW64mekjEjRw6gc" alt="" width="563"><figcaption><p>Get variables</p></figcaption></figure>
{% endtab %}

{% tab title="5. Delete" %}
{% hint style="info" %}

#### Delete

Delete is a **terminal** step. It does not pass rows downstream.\
It builds `DELETE` statements from the input stream.
{% endhint %}

{% hint style="danger" %}
Be careful with the comparators in this step.\
Always validate your criteria before you run.
{% endhint %}

1. Drag the Delete step onto the canvas.
2. Open the Delete properties dialog box.
3. Configure the database **Connection** and set **Table name** to `STG_SALES_DATA`.
4. Map stream fields to table fields for your delete criteria.

<figure><img src="/files/b14tz5ZHJVgVwhXVvRea" alt="" width="563"><figcaption><p>Delete step</p></figcaption></figure>

{% hint style="info" %}
This workshop uses criteria based on:

* `QUANTITYORDERED` and the `min_quantityordered` value
* `PRODUCTLINE` values from `productlines.csv`
  {% endhint %}
  {% endtab %}

{% tab title="6. Run and validate" %}
{% hint style="info" %}

#### Run and validate

Run the transformation, then validate the row counts and sample rows.
{% endhint %}

{% hint style="warning" %}
Re-run your baseline queries from the first tab.\
Confirm the results match your delete criteria.
{% endhint %}

1. Select **Run** in Spoon.
2. In your database tool, inspect `STG_SALES_DATA`.

<figure><img src="/files/3LXBQrx7nwO40j4J47pE" alt=""><figcaption><p>STG_SALES_DATA</p></figcaption></figure>
{% endtab %}
{% endtabs %}

<details>

<summary>Troubleshooting</summary>

**Nothing was deleted**\
Your criteria did not match any rows. Confirm your comparator and data types.

**Too many rows deleted**\
Your comparator is too broad, or you mapped the wrong field names.

**Delete step fails with type conversion errors**\
Cast `min_quantityordered` to a number with **Select values**.

</details>


# Data Cleansing

Traditional data cleansing techiques ..

{% hint style="info" %}
This workshop section focuses on demonstrating traditional data cleansing techniques using Pentaho Data Integration.

This dataset contains various issues:

* Duplicate records (John Doe, Alice Johnson)
* Inconsistent phone number formats
* Inconsistent date formats
* Missing values
* Inconsistent address formats
* Removing Duplicates

Your mission, should you wish to accept it .. is to build a workflow / pipeline that resolves the issues .. explain the decisions you have made and any suggest possible enhancements ..
{% endhint %}

```
CustomerID	FirstName	LastName	Email	                Phone	         BirthDate	Address
1	        John	        Doe	        john.doe@email.com	555-123-4567	 1985-03-15	123 Main St, City, CA, 12345
2	        Jane	        Smith	        jane.smith@email.com	(555) 987-6543	 03/22/1990	456 Elm Avenue, Town, AZ, 67890
3	        John	        Doe	        johnd@email.com	        5551234567	 1985-03-15	123 Main Street, City, CO, 12345
4	        Alice	        Johnson	        alice.j@email.com	555-555-5555	 1988-12-01	789 Oak Rd, Village, State, 54321
5	        Bob	        Williams	bob.w@email.com		                 1975-07-30	101 Pine Lane, Hamlet, State, 13579
6	        Emma	        Brown	        emma.brown@email.com	(555)246-8135	 05-19-1992	202 Cedar Blvd, Borough, NY, 24680
7	        Alice	        Johnson	        alice.johnson@email.com	555.555.5555	 12/01/1988	789 Oak Road, Village, FL, 54321
8	        Charlie	        Davis	        charlie.d@email.com	555-369-2587		        303 Maple Dr, City, State, 97531
9		                Taylor	        d.taylor@email.com	555-159-7532	 1982-09-25	404 Birch St, Town, State, 86420
10	        Grace	        Lee	        grace.lee@email.com	5557894561	 11-11-1995	505 Walnut Ave, City, State, 
...
```

{% hint style="info" %}
All the workshop files: ../Databases/CRUID/Workshop - Data Cleansing
{% endhint %}

{% tabs %}
{% tab title="1. Onboard Data" %}
{% hint style="info" %}
First step is to onboard the data .. currently its a CSV file - which could easily be onboarded - instead let's onboard into a table.

There's a number of databases already installed and configured, running as containers.
{% endhint %}

1. Log on to Portainer and check the MariaDB database container is up and running.

{% embed url="<https://localhost:9443/#!/2/docker/containers>" %}
Link to Docker Containers in Portainer
{% endembed %}

2. Execute the following script to create a sourceDB & targetDB databases.

'grant all' to pentaho\_user & pentaho\_admin with the password: 'password'.

```
CREATE DATABASE  IF NOT EXISTS sourceDB;
grant all on sourceDB.* to pentaho_user identified by 'password';
grant all on sourceDB.* to pentaho_admin identified by 'password';

USE sourceDB;

set session sql_mode=replace(@@sql_mode,'NO_ZERO_DATE','');
```

<figure><img src="/files/NZLPNCySNsKNMtln4xie" alt=""><figcaption><p>create sourceDB &#x26; targetDB</p></figcaption></figure>

***

**Transformation**

{% hint style="info" %}
Currently the customer records are in CSV format. It's optional, but, let's onboard into a database table as this is a more likely scenario.
{% endhint %}

<figure><img src="/files/eT7kmKmvzqR1uqDINIim" alt="" width="296"><figcaption><p>tr_onboard</p></figcaption></figure>

{% tabs %}
{% tab title="CSV File input" %}

1. Drag and drop a CSV File input onto the canvas.
2. Double-click to configure the following settings:

<figure><img src="/files/yC9EbNgfTJhPrIHmwZqR" alt=""><figcaption><p>CSV File input - customer_data.csv</p></figcaption></figure>

3. Increase the varchar (length) to prevent truncation.
4. After clicking on 'Get Fields', 'Preview' the data.

<figure><img src="/files/irIRzay0gtVIFek5gf9z" alt=""><figcaption><p>Preview data</p></figcaption></figure>
{% endtab %}

{% tab title="Table output" %}
{% hint style="info" %}
The MariaDB jdbc database driver has been copied to the /lib directory.
{% endhint %}

**Define Database Connection (MariaDB)**

1. In Spoon, click File > New > Transformation. Any one of these actions opens a new Transformation tab for you to begin designing your transformation:
   * By clicking New, then Transformation
   * By using the CTRL-N hot key
2. From within Spoon, Select:

File > New > Database Connection The Database Connection dialog box appears.

<figure><img src="/files/N756WJLwzYvAkNB7Gq9Y" alt=""><figcaption><p>MariaDB: sourceDB connection</p></figcaption></figure>

3. Enter the following details:

<table><thead><tr><th width="193">Section Name</th><th>Value</th></tr></thead><tbody><tr><td><strong>Connection Name</strong></td><td>MariaDB: sourceDB</td></tr><tr><td><strong>Connection Type</strong></td><td><mark style="color:red;"><strong>MySQL</strong></mark></td></tr><tr><td><strong>Host Name</strong></td><td>localhost</td></tr><tr><td><strong>Database Name</strong></td><td>sourceDB</td></tr><tr><td><strong>User name</strong></td><td>pentaho_admin</td></tr><tr><td><strong>Password</strong></td><td>password</td></tr></tbody></table>

4. Click Test.
5. Configure the Table ouput step with the following settings:

<figure><img src="/files/wvx52vGix0NoEOS1BMq9" alt="" width="473"><figcaption><p>create source_customer table</p></figcaption></figure>

6. Click on the SQL button & Execute.

<figure><img src="/files/0bOaEhRNMKZSlbKjcXvT" alt="" width="365"><figcaption><p>Execute SQL script</p></figcaption></figure>

7. Close all the windows & save / RUN transformation.

<figure><img src="/files/CNVdF1CyBFuxKpi3apkU" alt=""><figcaption><p>RUN Transformation</p></figcaption></figure>

8. Check the table has been created and populated with customer data.

<figure><img src="/files/O9lPHVdEw9nrhFoaVmmw" alt=""><figcaption><p>souurce_customer table</p></figcaption></figure>
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. Data Cleansing" %}
{% hint style="info" %}
Lets' breakdown the data cleansing workflows that need to be applied:

Resolve FirstName
{% endhint %}

x

x

x

{% tabs %}
{% tab title="FirstName" %}
{% hint style="info" %}
Let's start with resolving the FirstName.

Luckily we access to the National Address Database (NAD), which is a aggregation of data provided by state, local, and tribal governments.
{% endhint %}

<figure><img src="/files/RGxmGt8K9bUUtyLdJhfy" alt=""><figcaption><p>Resolve FirstName</p></figcaption></figure>

**If field value is null**

{% hint style="info" %}
The step "If field value is null" is able to replace nulls by a given value either by

1. Processing the complete row with all fields
2. Processing the complete row but only for specific field types (Number, String, Date etc.)
3. Processing the complete row but only for specific fields by name

Changing a numeric field type to an empty string field type can cause an error in the subsequent steps. The error will not appear until the next step accesses that field.
{% endhint %}

<figure><img src="/files/l57Lg1jBwjNyWMQct0ql" alt=""><figcaption><p>Replace NULL values</p></figcaption></figure>

***

**Filter FN**

{% hint style="info" %}
Replacing NULL values with 'Unknown' enables the FirstName field to be filtered.

You could also modify the Filter condition to: NOT
{% endhint %}

<figure><img src="/files/NXFW3lS2k4rdSgPoLNqE" alt=""><figcaption><p>Filter: FirstName = Unknown</p></figcaption></figure>

***

**Stream lookup FN**

{% hint style="info" %}
If the condition is TRUE, i.e FirstName = Unknown then send the row to Stream lookup FN.
{% endhint %}

<figure><img src="/files/2zsBTCeMaGixwd9nwZbo" alt=""><figcaption><p>Lookup FirstName and set as ResolvedFisrtName</p></figcaption></figure>

{% hint style="info" %}
Ok .. I know this is a training lab ..!

The Data grid (FisrtName) has customer data that could be coming from another system.

Fingers crossed the CustomerID references the same customer.

The key fields - hashed - ensure that the row returned 'matches' the row in the data stream.
{% endhint %}

***

Select values FN

{% hint style="info" %}
As we know the FirstName = Unknown, however, based on the lookup criteria its matched Resolved FirstName = David to the record. To 'map' the ResolvedFirstName to the FirstName data stream field, simply rename ..

So .. now when you merge the record back into the 'main' stream
{% endhint %}

<figure><img src="/files/nIgMAIBVP9OZVScW8FIW" alt=""><figcaption></figcaption></figure>

x

x
{% endtab %}

{% tab title="Second Tab" %}
x
{% endtab %}
{% endtabs %}

x
{% endtab %}
{% endtabs %}


# SCDs

Slowly Changing Dimensions ..

{% hint style="info" %}

#### **Overview**

Slowly Changing Dimensions (SCD) - dimensions that change slowly over time, rather than changing on regular schedule, time-base. In a Data Warehouse, there is a need to track changes in dimension attributes in order to report historical data. In other words, implementing one of the SCD types should enable users assigning proper dimension's attribute value for given date. Example of such dimensions could be: customer, geography, employee.

There are many approaches how to deal with SCD. The most popular are:

* Type 0 - The passive method
* Type 1 - Overwriting the old value
* Type 2 - Creating a new additional record
* Type 3 - Adding a new column
* Type 4 - Using historical table
* Type 6 - Combine approaches of types 1,2,3 (1+2+3=6)
  {% endhint %}

{% tabs %}
{% tab title="Type 0" %}
{% hint style="info" %}
**Type 0 - value does not change over time**

A type 0 slowly changing dimension is a dimension that never changes its attributes over time.

For example, the date of birth of a person is a type 0 attribute, because it does not change after it is recorded. A type 0 dimension can be used to store the original values of some attributes that are not relevant for historical analysis.
{% endhint %}
{% endtab %}

{% tab title="Type 1" %}
{% hint style="info" %}
**Type 1 - Overwrite with new value**

Overwriting the old value. In this method, no history of dimension changes is kept in the database. The old dimension value is simply overwritten with the new one. This type is easy to maintain and is often use for data which changes are caused by processing corrections (e.g. removal special characters, correcting spelling errors).
{% endhint %}

<figure><img src="/files/Y9TqJZWI6iu1yTPancmm" alt=""><figcaption><p>Type 1</p></figcaption></figure>
{% endtab %}

{% tab title="Type 2" %}
{% hint style="info" %}
**Type 2 - Historically track the value change**

Type 2 - Creating a new additional record. In this methodology, all history of dimension changes is kept in the database. You capture attribute change by adding a new row with a new surrogate key to the dimension table. Both the prior and new rows contain as attributes the natural key (or another durable identifier).

Also 'effective date' and 'current indicator' columns are used in this method. There could be only one record with current indicator set to 'Y'. For 'effective date' columns, i.e. start\_date and end\_date, the end\_date for current record usually is set to value 9999-12-31.

Introducing changes to the dimensional model in type 2 could be very expensive database operation so it is not recommended to use it in dimensions where a new attribute could be added in the future.
{% endhint %}

<figure><img src="/files/P3nLf4Puht32O5Rnhbl6" alt=""><figcaption><p>Type 2</p></figcaption></figure>
{% endtab %}

{% tab title="Type 3" %}
{% hint style="info" %}
**Type 3 - Add a column to track value changes**

Adding a new column. In this type, usually only the current and previous value of dimension is kept in the database. The new value is loaded into 'current/new' column and the old one into 'old/previous' column.

The history is limited to the number of column created for storing historical data. This is the least commonly needed technique.
{% endhint %}

<figure><img src="/files/CezTzkTLJHDADGI61JLz" alt=""><figcaption><p>Type 3</p></figcaption></figure>
{% endtab %}

{% tab title="Type 4" %}
{% hint style="info" %}
Using historical table. In this method, a separate historical table is used to track all dimension's attribute historical changes for each of the dimension. The 'main' dimension table keeps only the current data e.g. customer and customer\_history tables.
{% endhint %}

<figure><img src="/files/kDgINpQBfBAMI0iTjKyL" alt=""><figcaption><p>Type 4</p></figcaption></figure>
{% endtab %}

{% tab title="Type 6" %}
{% hint style="info" %}
Combine approaches of types 1,2,3 (1+2+3=6). In this type, we have in dimension table such additional columns as:\ <mark style="color:red;">current\_type</mark> - for keeping current value of the attribute. All history records for given item of attribute have the same current value.\ <mark style="color:red;">historical\_type</mark> - for keeping historical value of the attribute. All history records for given item of attribute could have different values.\ <mark style="color:red;">start\_date</mark> - for keeping start date of 'effective date' of attribute's history. <mark style="color:red;">end\_date</mark> - for keeping end date of 'effective date' of attribute's history. <mark style="color:red;">current\_flag</mark> - for keeping information about the most recent record. In this method to capture attribute change we add a new record as in type 2. The current\_type information is overwritten with the new one as in type 1. We store the history in a historical\_column as in type 3.
{% endhint %}

<figure><img src="/files/ANsIytWbTDVJCZ3Bw1j0" alt=""><figcaption><p>Type 6</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# SCDs

Slowly Changing Dimensions ..

{% hint style="warning" %}

#### Workshop - Slowly changing dimensions (SCDs)

Maintain a dimension table as attributes change over time.\
Use **Dimension Lookup/Update** for both **Type 1** and **Type 2** behavior.

**What you’ll do**

* Create a `DIM_SCD` table in `sampledata`
* Generate test rows with **Data Grid**
* Add timestamps with **Get System Info**
* Configure **Dimension Lookup/Update** for:
  * Lookup mode (read-only enrichment)
  * Update mode (writes)
* Test field strategies:
  * **Update** (Type 1 overwrite)
  * **Insert** (Type 2 new version)
  * **Punch through** (update all versions)

**Prerequisites**

* PDI installed and running
* A working MySQL connection to `sampledata`
* Familiarity with natural keys vs surrogate keys

**Estimated time:** 40 minutes
{% endhint %}

{% embed url="<https://www.loom.com/share/3ec2d5123814460d92085851c18daaee?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Dimension Lookup
{% endembed %}

***

{% hint style="info" %}
**Workshop files**

Download the following files.

Keep the filenames unchanged.

Save them in your workshop folder.
{% endhint %}

{% file src="/files/owfoon4eYQ2NYLncIdJz" %}

{% file src="/files/bkmeMXrhLGexaQdaJzAp" %}

***

{% hint style="info" %}
**Create a new transformation**

Use any of these options to open a new transformation tab:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

<figure><img src="/files/ifoV3zEDKa3rPve59CLp" alt=""><figcaption><p>tr_scd</p></figcaption></figure>

***

{% tabs %}
{% tab title="Transformation" %}
{% hint style="danger" %}

#### Transformation

Ensure the tr\_scd.ktr has been created before you explore the Workflows.
{% endhint %}

{% tabs %}
{% tab title="1. DIM\_SCD Table" %}
{% hint style="info" %}

#### **SCD table**

First step in these series of workshops, is to create a simple DIM\_SCD table that has the required fields to illustrate a Type 1 change - just overwrite the value - and Type 2 - where you need to record when any change in the value occurs.
{% endhint %}

1. Start PDI.

{% hint style="info" %}
{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

{% hint style="warning" %}
Create the Table in MySQL sampledata database.
{% endhint %}

2. In your DB management tool, execute the following statement.

```sql
CREATE TABLE `DIM_SCD` (
  `TK` bigint(10) NOT NULL,
  `version` int(11) DEFAULT '0',
  `id` int(11) DEFAULT '0',
  `city` tinytext,
  `date_from` datetime DEFAULT NULL,
  `date_to` datetime DEFAULT NULL,
  `last_update` datetime DEFAULT NULL,
  PRIMARY KEY (`TK`),
  KEY `idx_DIM_SCD_lookup` (`id`)
) ENGINE=InnoDB DEFAULT CHARSET=latin1;
```

<figure><img src="/files/eRDIitMlAXoFTsBMx0bu" alt=""><figcaption><p>DIM_SCD</p></figcaption></figure>
{% endtab %}

{% tab title="2. Data Grid" %}
{% hint style="info" %}

#### **Data grid**

This step allows you to enter a static list of rows in a grid. This is usually done for testing, reference or demo purposes.
{% endhint %}

1. Drag the Data Grid step onto the canvas.
2. Open the Data Grid properties dialog box.
3. Ensure the following details are configured, as outlined below:

<figure><img src="/files/9K3BEd8lLdmef8OxTLr9" alt=""><figcaption><p>Data grid - meta</p></figcaption></figure>

4. On the Data tab enter the following values:

<figure><img src="/files/bYbnAoLAyOtHx4pKYE25" alt=""><figcaption><p>Data grid - data</p></figcaption></figure>
{% endtab %}

{% tab title="3. Get System Info" %}
{% hint style="info" %}

#### **Get system info**

The Get System Info step retrieves information from the Kettle environment.

This step generates a single row with the fields containing the requested information. It also accepts input rows.
{% endhint %}

1. Drag the Get System Info step onto the canvas.
2. Open the Get System Info properties dialog box.
3. Ensure the following details are configured, as outlined below:

<figure><img src="/files/ByEffViXH23izWRnChKK" alt="" width="563"><figcaption><p>Get system info</p></figcaption></figure>
{% endtab %}

{% tab title="4. Dimension lookup/update" %}
{% hint style="info" %}

#### **Dimension lookup/update**

The Dimension Lookup/Update step allows you to implement Ralph Kimball's slowly changing dimension for both types: Type I (update) and Type II (insert) together with some additional functions. Not only can you use this step to update a dimension table, it may also be used to look up values in a dimension.
{% endhint %}

1. Drag the Dimension lookup/update step onto the canvas.
2. Open the Dimension lookup/update properties dialog box.
3. Ensure the following details are configured, as outlined below:

* **Connection:** your `sampledata` connection
* **Target table:** `DIM_SCD`
* **Technical key field:** `TK`
* **Version field:** `version`
* **Lookup key:** `id`

{% hint style="info" %}
You will switch between **Lookup mode** and **Update mode** in the workflows below.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Workflow 1 - Lookup" %}
{% hint style="info" %}
**Lookup**

This step is rarely used for lookups because specific steps exist. However, it does illustrate how referential integrity is maintained.
{% endhint %}

1. Add ‘London’ to the Data Grid.

<figure><img src="/files/NtNrv71jnc9gDBQkzFmy" alt=""><figcaption><p>Data grid - lookup</p></figcaption></figure>

{% hint style="warning" %}
To initially configure the fields set the step to Update mode.

Once completed remember to reset to lookup mode.
{% endhint %}

<figure><img src="/files/AGJUF1CuTZow9DPqIqhz" alt=""><figcaption><p>Configure fields</p></figcaption></figure>

2. Ensure that the Dimension Lookup / Update is set to: Lookup Mode.

<figure><img src="/files/zueP4vmMx2dqrvfKBG1P" alt=""><figcaption><p>set lookup mode</p></figcaption></figure>

{% hint style="info" %}
Sets the step to Update Mode with lookup keys: id

There’s no point adding the `last_update` field because we’re dealing with Type 1 changes.
{% endhint %}

3. Click on the Fields tab.

<figure><img src="/files/wCkyGNzEinXf4IBwEr6U" alt=""><figcaption><p>Set city to string</p></figcaption></figure>

***

{% hint style="info" %}
**Run the transformation**

Preview the output from **Dimension lookup/update**.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Preview data for the **Dimension lookup/update** step.

<figure><img src="/files/vEx3Me645Ry5t8TO6Z7S" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="info" %}
**Why does London have a TK 0 and a value of NULL?**

To maintain referential integrity, the record is assigned a TK 0, and as it doesn’t exist in the database, the value returned is NULL. As this is a lookup, no record is written to the database table.
{% endhint %}
{% endtab %}

{% tab title="Workflow 2 - Type 1" %}
{% hint style="info" %}
The Dimension Lookup/Update step allows you to implement Ralph Kimball's slowly changing dimension for both types:

**Type 1**

Overwriting the old value. In this method, no history of dimension changes is kept in the database. The old dimension value is simply overwritten with the new one. This type is easy to maintain and is often used for data where changes are caused by processing corrections (for example, removal of special characters, correcting spelling errors).
{% endhint %}

1. Open the Dimension Lookup / Update properties dialog box.
2. Ensure the following details are configured, as outlined below:

<figure><img src="/files/jCuS5DeZpVcdskJ0B7pv" alt=""><figcaption><p>Type 1</p></figcaption></figure>

3. Click on the Get Fields button.

{% hint style="info" %}
This will add the key fields used in the Lookup.

The keys that map dimension table rows to stream rows are: `id`
{% endhint %}

4. Click on the Fields tab and map the stream name field to the dimension name field and ensure:

<figure><img src="/files/gtUOKS0rJ5PNK3A2zTAo" alt=""><figcaption><p>Field mapping and set database strategy</p></figcaption></figure>

{% hint style="info" %}
There are several options available to insert / update the dimension record (row).

**Insert:**

This option implements a Type I & II slowly changing dimension policy. If the difference is detected for one or more mappings that have the Insert option, then a row is added to the dimension table.

**Update:**

This option simply updates the matched row. It can be used to implement a Type I slowly changing dimension.

**Punch through:**

The punch through option also performs an update. But instead of only updating the matched dimension row, it will update all versions of the row in a Type II slowly changing dimension.

**Date of last insert or update (without stream field as source):**

Use this option to let the step automatically maintain a date field that records the date of the insert or update using the system date field as source.

**Date of last insert (without stream field as source):**

Use this option to let the step automatically maintain a date field that records the date of the last insert using the system date field as source.

**Date of last update (without stream field as source):**

Use this option to let the step automatically maintain a date field that records the date of the last update using the system date field as source.

**Last version (without stream field as source):**

Use this option to let the step automatically maintain a flag that indicates if the row is the last version.
{% endhint %}

***

{% hint style="success" %}
**RUN Transformation**

At the moment it's all Type I changes .. Insert / Update record and Update when that record change occurred.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Click on the Preview tab:

<figure><img src="/files/ks7Ma0mfwHQMWntRE59w" alt=""><figcaption><p>Preview data</p></figcaption></figure>

3. Examine the table in your SQL Query Tool:

<figure><img src="/files/NKojqhem204zya3X8yiV" alt=""><figcaption><p>DIM_SCD - view data</p></figcaption></figure>

{% hint style="info" %}
**So what's happening behind TK 0?**

Kettle automatically inserts an additional record with a technical key of value 0 (for default or unknown values). This will only happen in the first execution. Below this record, you find the one record (London) from our sample dataset.

In update mode (update option is enabled) the step first performs a lookup of the dimension entry. The result of the lookup is different though. Not only the technical key is retrieved from the query, but also the dimension attribute fields. A field-by-field comparison then follows. The result can be one of the following situations:

* The record was not found, we insert a new record in the table.
* The record was found and any of the following is true:
* One or more attributes were different and had an "Insert" (Kimball Type II) setting: A new dimension record version is inserted.
* One or more attributes were different and had a "Punch through" (Kimball Type I) setting: These attributes in all the dimension record versions are updated.
* One or more attributes were different and had an "Update" setting: These attributes in the last dimension record version are updated.
* All the attributes (fields) were identical: No updates or insertions are performed.

If you mix Insert, Punch Through and Update options in this step, this algorithm acts like a Hybrid Slowly Changing Dimension. (it is no longer just Type I or II, it is a combination)
{% endhint %}

**Try different scenarios**

{% tabs %}
{% tab title="Insert / Update" %}

1. Double-click on the Data Grid step.
2. Add the following City: Madrid

<figure><img src="/files/OtzczAPgwaQ8lvqqis11" alt=""><figcaption><p>data Grid - Madrid</p></figcaption></figure>

3. Double-click on the Dimension lookup/update step.
4. Click on the Fields tab and set the following:

<figure><img src="/files/t90EBIyeGCtxFAUV4wPy" alt=""><figcaption><p>Field mapping and set database operation</p></figcaption></figure>

5. Execute the transformation and view the results.

<figure><img src="/files/55tg60JtYylWwrIuE95J" alt=""><figcaption><p>DIM_SCD update</p></figcaption></figure>

{% hint style="info" %}
Note that if the record does not exist (Madrid), then it gets inserted. Both records are updated with the same last\_update timestamp.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Workflow 3 - Type 2" %}
{% hint style="info" %}
The Dimension Lookup/Update step allows you to implement Ralph Kimball's slowly changing dimension for both types:

**Type 2**

Creating a new additional record. In this methodology, all history of dimension changes are kept in the database. You capture attribute change by adding a new row with a new surrogate key to the dimension table. Both the prior and new rows contain as attributes the natural key (or another durable identifier). Also 'effective date' and 'current indicator' columns are used in this method. There could be only one record with current indicator set to 'Y'. For 'effective date' columns, i.e. start\_date and end\_date, the end\_date for current record usually is set to value 9999-12-31. Introducing changes to the dimensional model in Type 2 could be very expensive database operation so it is not recommended to use it in dimensions where a new attribute could be added in the future.
{% endhint %}

1. Double-click on the Data Grid step.
2. Add the following City: Pariss (intentionally misspelled)

<figure><img src="/files/o5cwG8zWOwFs6xCZjCXR" alt=""><figcaption><p>Data grid - Pariss</p></figcaption></figure>

3. Keep the Fields strategy the same:

<figure><img src="/files/t90EBIyeGCtxFAUV4wPy" alt=""><figcaption><p>Field mapping and set database strategy</p></figcaption></figure>

**RUN Transformation**

{% hint style="success" %}
At the moment it's all Type I changes .. Insert / Update record and Update when that record change occurred.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Click on the Preview tab:

<figure><img src="/files/DYc5eReNz4miAcxQgavi" alt=""><figcaption><p>Preview data</p></figcaption></figure>

3. Examine the table in your SQL Query Tool:
4. Execute the transformation and view the results.

<figure><img src="/files/MAP2zxVbb56vdpOO9G5i" alt=""><figcaption><p>DIM_SCD - view data</p></figcaption></figure>

{% hint style="info" %}
As expected the Pariss record is inserted. Notice that the natural primary keys (id) are the same as the TK (Technical Key) and we're on version 1 for each record.
{% endhint %}

{% hint style="warning" %}
Let's now do a bit of data entry .. Correct the entry: Pariss to Paris
{% endhint %}

**Try out different Strategies**

{% hint style="info" %}
The workflows below will give you an idea of the different strategies that can be implemented.
{% endhint %}

{% tabs %}
{% tab title="Type 2" %}
{% hint style="info" %}
Any records that are changed are:

* Archived with date\_from to date\_to timestamp.
* Record is assigned Version 1. New record Version 2
* Note the Keys.
  {% endhint %}

1. Double-click on the Data Grid step.
2. Edit the Pariss value to: Paris

<figure><img src="/files/AevxxOrxXXJPyqZh4533" alt=""><figcaption><p>Data grid - edit Paris</p></figcaption></figure>

3. Double-click on the Dimension lookup/update step.
4. Click on the Fields tab and set the following:

{% hint style="warning" %}
The reason for selecting Insert is obvious.. you need to insert a new record that tracks the change. If you select Update, then a Type 1 change occurs; i.e. the original record is updated.
{% endhint %}

<figure><img src="/files/DBqjtNCkkhG4aPhNgGF9" alt=""><figcaption><p>Field mapping and set database strategy</p></figcaption></figure>

**RUN the Transformation**

1. Click the Run button in the Canvas Toolbar.
2. Examine the table in your SQL Query Tool:.

<figure><img src="/files/2wk6CwSh3TQIUilICTZH" alt=""><figcaption><p>DIM_SCD - view data</p></figcaption></figure>

{% hint style="info" %}

* A new record has been inserted, with the value ‘Paris’, with an updated date\_from and last\_update timestamp, and version.
* Notice that the `last_update` field has also been updated for the other records.
  {% endhint %}
  {% endtab %}

{% tab title="Punch through" %}
{% hint style="info" %}
**Punch through:**

The punch through option also performs an update. But instead of only updating the matched dimension row, it will update all versions of the row in a Type II slowly changing dimension.

So let's update all versions to Paris.
{% endhint %}

1. Let's start with a clean table to clearly illustrate the results. Truncate the table.

```sql
TRUNCATE TABLE DIM_SCD;
```

2. Repeat the workflow above and check the results.

{% hint style="warning" %}

* Ensure you have set 'Pariss' as the value for city in the data grid step.
* Set update as your initial strategy.
  {% endhint %}

<figure><img src="/files/MAP2zxVbb56vdpOO9G5i" alt=""><figcaption><p>Preview data</p></figcaption></figure>

3. Double-click on the Data Grid step.
4. Edit the Pariss value to: Paris

   <figure><img src="/files/AevxxOrxXXJPyqZh4533" alt=""><figcaption><p>Data grid - edit Paris</p></figcaption></figure>
5. Double-click on the Dimension lookup/update step.
6. Click on the Fields tab and set the following:

<figure><img src="/files/INnngsjQYSB3KJKo7pmR" alt=""><figcaption><p>Field mapping and set database strategy</p></figcaption></figure>

**RUN the Transformation**

1. Click the Run button in the Canvas Toolbar.
2. Examine the table in your SQL Query Tool:

<figure><img src="/files/2wk6CwSh3TQIUilICTZH" alt=""><figcaption><p>DIM_SCD - view data</p></figcaption></figure>

**Punch through**

1. Edit the Fields tab in the Dimension Lookup/update step to Punch through:

<figure><img src="/files/MIgbIop4bCRwbaywkduZ" alt=""><figcaption><p>Field mapping and set database strategy</p></figcaption></figure>

**RUN the Transformation**

1. Click the Run button in the Canvas Toolbar.
2. Examine the table in your SQL Query Tool:

<figure><img src="/files/EkBzHUEBhLZDSDQOUWmP" alt=""><figcaption><p>DIM_SCD - view data</p></figcaption></figure>

{% hint style="info" %}
What happened

Remember a Punch through updates the fields where the records match ..

* All the records have a last\_update field .. so this will be updated
* As both our Pariss & Paris records have a matching key id=3 then the archived record will be updated with the current city value .. Paris
  {% endhint %}
  {% endtab %}
  {% endtabs %}
  {% endtab %}
  {% endtabs %}


# Storage

{% hint style="info" %}

#### **Overview**

**Block Storage** operates at the raw block level using high-performance protocols like Fibre Channel (FC) at 8/16/32 Gbps or iSCSI over Ethernet, providing direct-attached storage or SAN connectivity with sub-millisecond latency. It presents logical unit numbers (LUNs) as raw disk volumes to the operating system, making it optimal for databases, virtual machine disk images, and transactional workloads requiring consistent IOPS performance and low-level disk control.

**File Storage** operates at the file system layer using network protocols like NFS (supporting NFSv3/v4 with features like client-side caching and Kerberos authentication) or SMB/CIFS (with opportunistic locking and distributed file system capabilities), providing POSIX-compliant file semantics with hierarchical directory structures, metadata management, and concurrent access controls ideal for content repositories and shared application data.

**Object Storage** uses RESTful HTTP/HTTPS APIs over TCP/IP, storing data as objects with unique identifiers in flat namespaces within buckets or containers, implementing eventual consistency models and offering features like versioning, lifecycle policies, cross-region replication, and virtually unlimited horizontal scaling through distributed hash tables, making it perfect for cloud-native applications, content distribution, backup repositories, and big data analytics requiring petabyte-scale storage with global accessibility.

MinIO is an open-source object storage solution that's compatible with Amazon S3's API. It's particularly popular for private cloud deployments and can be run on-premises or in any cloud environment. MinIO excels at high-performance workloads and is often used in conjunction with Kubernetes for scalable container deployments.

Amazon S3 (Simple Storage Service) is the industry standard for cloud object storage, offering virtually unlimited scalability, 99.999999999% durability, and extensive integration with AWS services. It provides different storage tiers (like Standard, Infrequent Access, and Glacier) to optimize costs based on access patterns.

Pentaho Content Platform (HCP) is an enterprise-grade object storage system that focuses on data governance, compliance, and security. It offers advanced features like data classification, retention policies, and WORM (Write Once, Read Many) capabilities. HCP can be deployed on-premises or in hybrid cloud configurations and supports multiple protocols including S3 compatibility.
{% endhint %}

<figure><img src="/files/DkjulDZjSHQ4PJGwpwvu" alt=""><figcaption></figcaption></figure>

***


# MinIO

Access S3 type Object Store - VFS (Virtual File System) ..

{% hint style="info" %}

#### **MinIO**

MinIO is a high-performance, Kubernetes-native object storage system designed for cloud-native applications. Built from the ground up to be compatible with Amazon S3, MinIO offers a lightweight yet powerful alternative for organizations looking to deploy object storage in their own infrastructure.

At its core, MinIO provides distributed object storage with performance characteristics. It's capable of handling millions of operations per second and can store petabytes of data while maintaining sub-millisecond latency. This performance is achieved through a simplified architecture that eliminates complex dependencies and optimizes for modern hardware capabilities.

One of MinIO's key strengths lies in its versatility. It can be deployed virtually anywhere - from bare metal servers to public, private, and edge cloud environments. Organizations particularly value its seamless integration with Kubernetes, making it an ideal choice for containerized environments. MinIO's open-source nature also provides transparency and flexibility that many enterprises require for their data infrastructure needs.
{% endhint %}

<figure><img src="/files/1pgsPzESBCaBvC9u08r8" alt=""><figcaption><p>MinIO</p></figcaption></figure>

{% embed url="<https://min.io/>" %}
Link to MinIO
{% endembed %}

***


# MinIO

Hands-on workshops using MinIO as an S3-compatible object store.

{% hint style="warning" %}

#### Workshop series: PDI + MinIO (S3)

Build hands-on Pentaho Data Integration (PDI) transformations that read from and write to **MinIO** using **VFS**.

Workshops get harder as you go. Start with CSV joins. Then move into XML/JSON parsing, reconciliation, and multi-format ingestion.

**Workshops in this series**

* Sales Dashboard (CSV inputs + lookups + output)
* Inventory Reconciliation (XML + CSV + variance detection)
* Customer 360 (multi-source joins + metrics)
* Clickstream Funnel (sessionization + pivoting)
* Log Parsing (regex + time-series checks)
* Data Lake Ingestion (schema normalization + validation)

**You’ll practice**

* Connecting to MinIO buckets with VFS
* Reading and writing objects with `pvfs://MinIO/...` paths
* Joining and enriching streams (lookups and joins)
* Parsing XML and JSON
* Validating and shaping data for a curated layer

**Prerequisites:** MinIO running with sample data populated; basic transformation concepts; basic joins and aggregations

**Estimated time:** 4–6 hours total (each workshop is \~20–60 minutes)
{% endhint %}

<table><thead><tr><th width="187">Workshop</th><th>Key Skills</th></tr></thead><tbody><tr><td>Sales Dashboard</td><td>joins, lookups, aggregations</td></tr><tr><td>Inventory Reconciliation</td><td>XML parsing, outer joins, variance</td></tr><tr><td>Customer 360</td><td>multi-source, JSONL, calculations</td></tr><tr><td>Clickstream Funnel</td><td>sessionization, pivoting</td></tr><tr><td>Log Parsing</td><td>regex, time-series analysis</td></tr><tr><td>Data Lake Ingestion</td><td>schema normalization, validation</td></tr></tbody></table>

{% hint style="danger" %}
Complete the setup first: [Storage: MinIO](/pentaho-data-integration/setup/data-sources/storage#minio)
{% endhint %}

1. Verify that MinIO is running and populated.

```bash
# Check MinIO is running
curl -sf http://localhost:9000/minio/health/live && echo "MinIO OK" || echo "MinIO not running"

# Verify data exists (using mc client)
mc ls minio-local/raw-data --recursive
```

2. Start Pentaho Data Integration.

{% hint style="info" %}
Start Pentaho Data Integration (Spoon).

{% tabs %}
{% tab title="Windows (PowerShell)" %}

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endtab %}

{% tab title="macOS / Linux" %}

```bash
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endtab %}
{% endtabs %}
{% endhint %}

***

{% tabs %}
{% tab title="Sales Dashboard" %}
{% hint style="warning" %}

#### Sales Dashboard

The workshop demonstrates how Pentaho Data Integration enables organizations to rapidly create denormalized fact tables that power real-time business intelligence dashboards. By integrating data from multiple sources (customer data, product catalogs, and sales transactions), business users gain immediate access to actionable insights without waiting for IT to build complex data warehouses.

**Scenario:** A mid-sized e-commerce company needs to track daily sales performance across products, customer segments, and regions. Currently, sales managers wait 24-48 hours for IT to generate reports from disparate systems. With PDI, they can automate this process and refresh dashboards hourly.

**Key Stakeholders:**

* Sales Directors: Need to identify top-performing products and regions
* Marketing Teams: Require customer segmentation for targeted campaigns
* Finance: Need accurate revenue reporting by product category
* Operations: Must monitor inventory turnover rates
  {% endhint %}

***

{% hint style="info" %}
**Workshop files**

These files are already in MinIO:

* `pvfs://MinIO/raw-data/csv/sales.csv`
* `pvfs://MinIO/raw-data/csv/products.csv`
* `pvfs://MinIO/raw-data/csv/customers.csv`

Output path used later: `pvfs://MinIO/staging/dashboard/`
{% endhint %}

{% file src="/files/qAmiEhBdNzKu5ugKPvsM" %}

***

<figure><img src="/files/z8GQ0foavHg7YOap5J9C" alt=""><figcaption><p>Sales Dashboard</p></figcaption></figure>

{% hint style="info" %}
Create a new transformation.

Use any of these options:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

Follow the steps to create the transformation:

{% tabs %}
{% tab title="1. Read Data Sources" %}
{% hint style="info" %}

#### **Text File Input**

The Text File Input step is used to read data from a variety of different text-file types. The most commonly used formats include Comma Separated Values (CSV files) generated by spreadsheets and fixed width flat files.

The Text File Input step provides you with the ability to specify a list of files to read, or a list of directories with wild cards in the form of regular expressions. In addition, you can accept filenames from a previous step making filename handling more even more generic.
{% endhint %}

<figure><img src="/files/l0loAmeYXrST39llxG2g" alt=""><figcaption><p>Text file inputs</p></figcaption></figure>

{% hint style="info" %}
VFS connection names are case-sensitive. These examples assume your connection name is `MinIO`.
{% endhint %}

1. Drag & drop 3 Text File Input Steps onto the canvas.
2. Save transformation as: `sales_dashboard_etl.ktr` in your workshop folder.

***

**Sales (Order Management)**

1. Double-click on the first TFI step, and configure with the following properties:

| Setting          | Value                                 |
| ---------------- | ------------------------------------- |
| Step name        | `Sales`                               |
| Filename         | `pvfs://MinIO/raw-data/csv/sales.csv` |
| Delimiter        | ,                                     |
| Head row present | ✅                                     |
| Format           | mixed                                 |

<figure><img src="/files/XlXKeRhImBeHbuaSmdUZ" alt=""><figcaption><p>Select - sales.csv from VFS connections</p></figcaption></figure>

2. Click: **Get Fields** to auto-detect columns.

{% hint style="info" %}
**Business Logic:** Note that `sale_amount` may differ from `price * quantity` due to:

* Volume discounts
* Promotional pricing
* Customer-specific pricing tiers
* Currency conversion (for international sales)
  {% endhint %}

<figure><img src="/files/nLGRIikQYnpXPcuFjQy7" alt=""><figcaption><p>Get Fields - Sales</p></figcaption></figure>

3. Preview data.

<figure><img src="/files/RNrGgMfChgJ6FwVj3Qzh" alt=""><figcaption><p>Preview data - Sales</p></figcaption></figure>

{% hint style="info" %}
**Business Significance:**

* `sale_amount`: Actual revenue (may include discounts)
* `quantity`: Volume metrics for demand planning
* `payment_method`: Payment preference insights
* `status`: Filter out cancelled/refunded orders
  {% endhint %}

***

**Products (ERP system)**

1. Double-click on the second TFI step, and configure with the following properties:

<table><thead><tr><th width="165.5">Setting</th><th>Value</th></tr></thead><tbody><tr><td>Step name</td><td><code>Products</code></td></tr><tr><td>Filename</td><td><code>pvfs://MinIO/raw-data/csv/products.csv</code></td></tr><tr><td>Delimiter</td><td>,</td></tr><tr><td>Head row present</td><td>✅</td></tr><tr><td>Format</td><td>mixed</td></tr></tbody></table>

<figure><img src="/files/NU4Daf3mNASlhknTwnz8" alt=""><figcaption><p>Select - products.csv from VFS connections</p></figcaption></figure>

2. Click: **Get Fields** to auto-detect columns.

<figure><img src="/files/AxuVByWQb1QbFvSv50AD" alt=""><figcaption><p>Get Fields - Customers</p></figcaption></figure>

3. Preview the data.

<figure><img src="/files/hn3ULfxvGNDbqt8SbzdO" alt=""><figcaption><p>Preview data - Products</p></figcaption></figure>

{% hint style="info" %}
**Business Significance:**

* `category`: Enables product performance analysis by segment
* `price`: Base pricing for margin calculations
* `stock_quantity`: Inventory turnover insights
  {% endhint %}

***

**Customers (CRM System)**

1. Double-click on the third TFI step, and configure with the following properties:

<table><thead><tr><th width="186">Setting</th><th>Value</th></tr></thead><tbody><tr><td>Step name</td><td><code>Customers</code></td></tr><tr><td>Filename</td><td><code>pvfs://MinIO/raw-data/csv/customers.csv</code></td></tr><tr><td>Delimiter</td><td>,</td></tr><tr><td>Header row present</td><td>✅</td></tr><tr><td>Format</td><td>mixed</td></tr></tbody></table>

<figure><img src="/files/qlGSSjv4O4UDRABGcAcF" alt=""><figcaption><p>Select - customers.csv from VFS connections</p></figcaption></figure>

2. Click: **Get Fields** to auto-detect columns.

<figure><img src="/files/VFuUGVY6mL6gkeYXfpSF" alt=""><figcaption><p>Get Fields - Customers</p></figcaption></figure>

3. Preview the data.

<figure><img src="/files/no1qconeb8vVB9MwrpJ5" alt=""><figcaption><p>Preview data - Customers</p></figcaption></figure>

{% hint style="info" %}
**Business Significance:**

* `customer_id`: Primary key for joining to sales
* `country`: Critical for geographic segmentation
* `status`: Identifies churned vs. active customers
* `registration_date`: Enables customer tenure analysis
  {% endhint %}
  {% endtab %}

{% tab title="2. Stream Lookup" %}
{% hint style="info" %}

#### Stream Lookup

A **Stream lookup** step enriches rows by looking up matching values from another stream.

In a transformation, you feed your main rows into one hop and a reference dataset into the other hop. The step then matches rows using key fields and returns the lookup fields on the output. It’s the in-memory alternative to a database lookup, but the reference stream must be available in the same transformation flow.
{% endhint %}

<figure><img src="/files/kymRfUl1AELaQtUpydTz" alt=""><figcaption><p>Lookups</p></figcaption></figure>

1. Drag & drop 2 **Stream lookup** steps onto the canvas.
2. Save transformation as: `sales_dashboard_etl.ktr` in your workshop folder.

***

**Product Lookup**

1. Draw a hop between **Sales** and **Product Lookup**.
2. Draw a hop between **Products** and **Product Lookup**.

{% hint style="info" %}
The Sales is acting as our Fact table. It holds the transaction data for our Products & Customers.
{% endhint %}

3. Double-click on the 'Product Lookup' step, and configure with the following properties:

| Tab     | Setting               | Value            |
| ------- | --------------------- | ---------------- |
| General | Step name             | `Product Lookup` |
| General | Lookup step           | `Products`       |
| Keys    | Field (from Sales)    | `product_id`     |
| Keys    | Field (from Products) | `product_id`     |

4. In **Values to retrieve**, add:
   * `product_name` (rename to `product_name`)
   * `category` (rename to `product_category`)
   * `price` (rename to `unit_price`)

<figure><img src="/files/BiSRxM3nbWtPXFmNnQsX" alt=""><figcaption><p>Product Lookup</p></figcaption></figure>

***

**Customers Lookup**

1. Draw a hop between **Product Lookup** and **Customers Lookup**.
2. Draw a hop between **Customers** and **Customers Lookup**.
3. Double-click **Customers Lookup**, and configure the following properties:

| Setting            | Value              |
| ------------------ | ------------------ |
| Step name          | `Customers Lookup` |
| Lookup step        | `Customers`        |
| Key field (stream) | `customer_id`      |
| Key field (lookup) | `customer_id`      |

4. Values to retrieve:
   * `first_name`
   * `last_name`
   * `country` (rename to `customer_country`)
   * `status` (rename to `customer_status`)

<figure><img src="/files/BsDF1EkIBHhICaeD4Ef6" alt=""><figcaption><p>Customers Lookup</p></figcaption></figure>

***

**Preview data**

1. Save the transformation.
2. RUN & Preview the data.

<figure><img src="/files/XM6VmDznx6e0n3iJP7qR" alt=""><figcaption><p>Lookups - Preview data</p></figcaption></figure>
{% endtab %}

{% tab title="3. Calculator" %}
{% hint style="info" %}

#### Calculator

The Calculator step provides predefined functions that you can run on input field values. Use Calculator as a quick alternative to custom JavaScript for common calculations.

To use Calculator, specify the input fields and the calculation type, and then write results to new fields. You can also remove temporary fields from the output after all values are calculated.
{% endhint %}

<figure><img src="/files/kcC0MNGKYV05apB6LLlr" alt=""><figcaption><p>Calculator step</p></figcaption></figure>

1. Drag & drop a 'Calculator' step onto the canvas.
2. Draw a Hop from the 'Customers Lookup' step to the 'Calculator' step.
3. Double-click on the 'Calculator' step, and configure the following properties:

<table><thead><tr><th width="161">New field</th><th>Calculation</th><th>Field A</th><th>Field B</th><th>Value type</th></tr></thead><tbody><tr><td><code>line_total</code></td><td>A * B</td><td>quantity</td><td>unit_price</td><td>Number</td></tr><tr><td><code>discount_amount</code></td><td>A - B</td><td>sale_amount</td><td>line_total</td><td>Number</td></tr></tbody></table>

<figure><img src="/files/gmhpU8D2UrK42NKhNW7J" alt=""><figcaption><p>Calculator</p></figcaption></figure>

***

**Preview data**

1. Save the transformation.
2. RUN & Preview the data.

<figure><img src="/files/xixt3MvnDObaRMNKsXyk" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="info" %}
**Business Insight Enabled:**

* **Positive `discount_amount`:** Customer received a discount (common)
* **Negative `discount_amount`:** Customer paid more than list price (expedite, premium, etc.)
* **Zero `discount_amount`:** Sold at list price
  {% endhint %}
  {% endtab %}

{% tab title="4. Formula" %}
{% hint style="info" %}

#### Formula

The Formula step can calculate Formula Expressions within a data stream. It can be used to create simple calculations like \[A]+\[B] or more complex business logic with a lot of nested if / then logic.
{% endhint %}

<figure><img src="/files/x9jwLThDa8iTJKxTRrqk" alt=""><figcaption><p>Formula step</p></figcaption></figure>

1. Drag & drop a 'Formula' step onto the canvas.
2. Draw a Hop from the 'Calculator' step to the 'Formula' step.
3. Double-click on the 'Formula' step, and configure the following properties:

<table><thead><tr><th width="190">New Field</th><th>Formula</th></tr></thead><tbody><tr><td>customer_full_name</td><td>CONCATENATE([first_name];" ";[last_name])</td></tr><tr><td>is_high_value</td><td>IF([sale_amount]>500;"Yes";"No")</td></tr></tbody></table>

<figure><img src="/files/IDPxHR8DMB4M7okF6lLM" alt=""><figcaption><p>Formula step</p></figcaption></figure>

***

**Preview data**

1. Save the transformation.
2. RUN & Preview the data.

<figure><img src="/files/WJcrQDdi4mTrQb6Gzz8t" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="info" %}
**Business Applications:**

* **is\_high\_value:** Trigger VIP customer service workflows
  {% endhint %}
  {% endtab %}

{% tab title="5. Add Constants" %}
{% hint style="info" %}

#### Add Constants

The Add constant values step is a simple and high performance way to add constant values to the stream.
{% endhint %}

<figure><img src="/files/MSY8dJUow1gXzlSSRWTK" alt=""><figcaption><p>Add constants</p></figcaption></figure>

1. Drag & drop 'Add constants' step onto the canvas.
2. Draw a Hop from the 'Formula' step to the 'Add constants ' step.
3. Double-click on the 'Add constants' step, and configure the following properties:

| Name          | Type   | Value            |
| ------------- | ------ | ---------------- |
| `data_source` | String | `minio_workshop` |

<figure><img src="/files/KGpyhLd8Ie1QzPKkKVQ6" alt=""><figcaption><p>Add constants</p></figcaption></figure>
{% endtab %}

{% tab title="6. Get System info" %}
{% hint style="info" %}

#### Get system info

This step retrieves system information from the Kettle environment. The step includes a table where you can designate a name and assign it to any available system info type you want to retrieve. This step generates a single row with the fields containing the requested information.

It can also accept any number of input streams, aggregate any fields defined by this step, and send the combined results to the output stream.
{% endhint %}

<figure><img src="/files/WjxXrhIllYmI5kykNkrw" alt=""><figcaption><p>get system info</p></figcaption></figure>

1. Drag & drop 'Get system info' step onto the canvas.
2. Draw a Hop from the 'Add constants' step to the 'Get system info ' step.
3. Double-click on the **Get system info** step, and configure the following properties:

| Name           | Type                   |
| -------------- | ---------------------- |
| etl\_timestamp | system date (variable) |

<figure><img src="/files/NUKwwczMuewFlAqcKASJ" alt=""><figcaption><p>Get system info</p></figcaption></figure>
{% endtab %}

{% tab title="7. Select Values" %}
{% hint style="info" %}

#### **Select Values**

The Select Values step can perform all the following actions on fields in the PDI stream:

**Select fields** - The Select Values step can perform all the following actions on fields in the PDI stream.

**Remove fields** - Use this tab to remove fields from the input stream.

**Meta-data** - Use this tab to change field types, lengths, and formats.
{% endhint %}

<figure><img src="/files/pRqzefMgYCnDit47qlS6" alt=""><figcaption><p>Select values</p></figcaption></figure>

1. Drag & drop a 'Select values' step onto the canvas.
2. Draw a Hop from the 'Get system info' step to the 'Select values' step.
3. Double-click on the 'Select values' step, and configure the following properties:
4. On **Select & Alter** tab, choose fields in order:

* sale\_id
* sale\_date
* customer\_id
* customer\_full\_name
* customer\_country
* customer\_status
* product\_id
* product\_name
* product\_category
* quantity
* unit\_price
* sale\_amount
* line\_total
* discount\_amount
* is\_high\_value
* payment\_method
* status (rename to `sale_status`)
* etl\_timestamp
* data\_source

<figure><img src="/files/pp1Byq5JvAmtFst7wlkI" alt=""><figcaption><p>Select</p></figcaption></figure>

***

**Preview data**

1. Save the transformation.
2. RUN & Preview the data.

<figure><img src="/files/hslHulDBTUZJQhDOADGn" alt=""><figcaption><p>Preview data</p></figcaption></figure>
{% endtab %}

{% tab title="8. Text File Output" %}
{% hint style="info" %}

#### Text file output

The Text File Output step exports rows to a text file.

This step is commonly used to generate delimited files (for example, CSV) that can be read by spreadsheet applications, and it can also generate fixed-length output.

You can’t run this step in parallel to write to the same file.

If you need to run multiple copies, select Include stepnr in filename and merge the resulting files afterward.
{% endhint %}

<figure><img src="/files/91XvtGEsAWMXwgWZGCOm" alt=""><figcaption><p>Text File output</p></figcaption></figure>

1. Drag & drop a **Text file output** step onto the canvas.
2. Draw a Hop from the 'Select values' step to the 'Write to staging' step.
3. Double-click on the 'Write to staging' step, and configure with the following properties:

| Setting                       | Value                                       |
| ----------------------------- | ------------------------------------------- |
| Step name                     | `Write to Staging`                          |
| Filename                      | `pvfs://MinIO/staging/dashboard/sales_fact` |
| Extension                     | `csv`                                       |
| Include date/time in filename | ✅                                           |
| Separator                     | ,                                           |
| Add header                    | ✅                                           |

{% hint style="warning" %}
Select **Get fields** to populate the output fields.
{% endhint %}

{% hint style="info" %}
**Business Benefit:** Timestamped files enable:

* **Historical tracking:** "What did the data look like last Tuesday?"
* **Incremental processing:** Keep processing latest file without overwriting history
* **Rollback capability:** "The 3pm run had bad data, revert to 2pm version"
  {% endhint %}

***

**MinIO**

1. Save the transformation.
2. Log into MinIO:

<figure><img src="/files/whbPHxeIYRXF1r7KyWhh" alt=""><figcaption><p>MinIO - Dashboard data</p></figcaption></figure>

***

**Checklist**

* [ ] Three Text file inputs configured (reading CSV from S3)
* [ ] Product lookup working (no null product names)
* [ ] Customer lookup working (no null countries)
* [ ] Calculations producing correct values
* [ ] Fields in correct order
* [ ] Output file created in staging bucket
* [ ] All 15 sales records processed
  {% endtab %}
  {% endtabs %}
  {% endtab %}

{% tab title="Inventory Reconciliation" %}
{% hint style="warning" %}

#### Inventory Reconciliation - XML + CSV Integration

This workshop demonstrates how Pentaho Data Integration eliminates costly inventory discrepancies by automatically reconciling data between warehouse management systems (XML feeds) and ERP product catalogs (CSV files). Organizations lose millions annually due to inventory inaccuracies, stockouts, and overstocking. PDI's ability to parse complex XML and perform full outer joins enables real-time discrepancy detection that would require hours of manual spreadsheet work.

**Business Value Delivered:**

* **Cost Reduction:** Eliminate manual reconciliation labor ($75K-150K annually per analyst)
* **Inventory Optimization:** Reduce excess inventory carrying costs by 15-25%
* **Stockout Prevention:** Identify missing items before customers notice
* **Compliance:** Audit trail for SOX, ISO 9001, and supply chain regulations
* **Real-Time Visibility:** Know your actual inventory position within minutes, not days

**Scenario:** A manufacturing company operates 12 distribution warehouses. Each warehouse uses a legacy WMS (Warehouse Management System) that exports XML inventory files nightly. The corporate ERP system maintains a CSV product master catalog. Discrepancies cause:

* **Phantom stock:** ERP shows item in stock, warehouse says it's not → Lost sales
* **Ghost inventory:** Warehouse has items ERP doesn't recognize → Dead capital
* **Quantity variances:** Mismatches of 10+ units trigger expensive physical counts

**Key Stakeholders:**

* **Supply Chain Directors:** Need accurate inventory positions across all locations
* **Warehouse Managers:** Require daily reconciliation reports to prioritize cycle counts
* **Finance Teams:** Must report accurate inventory valuations for financial statements
* **Procurement:** Need to identify slow-moving items and prevent overstocking
  {% endhint %}

***

{% hint style="info" %}
**Workshop files**

These files are already in MinIO:

* `pvfs://MinIO/raw-data/xml/inventory.xml`
* `pvfs://MinIO/raw-data/csv/products.csv`

Planned output path: `pvfs://MinIO/staging/inventory/reconciliation/`
{% endhint %}

<figure><img src="/files/iMGHIMSukH4yCfdIN32J" alt=""><figcaption><p>Inventory reconciliation</p></figcaption></figure>

{% hint style="info" %}
Create a new transformation.

Use any of these options:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

***

Follow the steps to create the transformation:

{% tabs %}
{% tab title="1. Data Source streams" %}
{% tabs %}
{% tab title="1. Read Warehouse" %}
{% hint style="info" %}

#### Get data from XML

{% endhint %}

1. Drag & drop 'Get data from XML' onto the canvas.
2. Save transformation as: `inventory_reconciliation.ktr` in your workshop folder.
3. Double-click on the 'Get data from XML' step, and configure with the following properties:

<table><thead><tr><th width="186">Setting</th><th>Value</th></tr></thead><tbody><tr><td>Step name</td><td>Read Warehouse XML</td></tr><tr><td>File or directory</td><td><code>pvfs://MinIO/raw-data/xml/inventory.xml</code></td></tr><tr><td>Loop XPath</td><td><code>/inventory/items/item</code></td></tr><tr><td>Encoding</td><td><code>UTF-8</code></td></tr><tr><td>Ignore comments</td><td>✅</td></tr><tr><td>Validate XML</td><td>No</td></tr><tr><td>Ignore empty file</td><td>✅</td></tr></tbody></table>

{% hint style="info" %}
**XPath Explanation:**

* `/inventory` = Start at root element
* `/items` = Navigate to items container
* `/item` = Loop over each item element
  {% endhint %}

4. Browse & Add the path to the inventory.xml
5. Click on the Content tab

<figure><img src="/files/kCiQo09ujNftYaSxbWM2" alt=""><figcaption><p>Configure XPath</p></figcaption></figure>

6. Click on the Fields tab & Get Fields.
7. Remap the fields & Preview rows.

{% hint style="info" %}
**Business Field Naming:**

* Prefix with `warehouse_` to distinguish from ERP fields later
* `warehouse_quantity` vs. `stock_quantity` makes joins clearer
* Keep original field names in a data dictionary for auditing
  {% endhint %}

| Name                  | XPath         |
| --------------------- | ------------- |
| warehouse\_item\_name | name          |
| warehouse\_quantity   | quantity      |
| warehouse\_location   | location      |
| last\_physical\_count | last\_checked |

<figure><img src="/files/WCmT4hfNb5jGeyUaxRWh" alt=""><figcaption><p>Remap field names &#x26; Preview data</p></figcaption></figure>

{% hint style="info" %}
Next: configure the product catalog input, then join the two streams.
{% endhint %}
{% endtab %}

{% tab title="2. Read Product Catalog" %}
{% hint style="info" %}
Status: **Draft**. Add a **Text file input** step for `pvfs://MinIO/raw-data/csv/products.csv`.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. Join" %}
{% hint style="info" %}
Status: **Draft**. Join warehouse items to the ERP product master using a full outer join.
{% endhint %}
{% endtab %}

{% tab title="3. Output" %}
{% hint style="info" %}
Status: **Draft**. Write discrepancy rows to `pvfs://MinIO/staging/inventory/reconciliation/`.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Customer 360" %}
{% hint style="warning" %}

#### Customer 360

Create unified customer profiles combining demographic data, purchase history, and behavioral events.

**Skills:** Multiple joins, JSONL parsing, aggregations, calculated metrics
{% endhint %}

<figure><img src="/files/RF3rEFvMNBDxwjiKGcuB" alt=""><figcaption><p>Customer 360</p></figcaption></figure>

{% hint style="info" %}
**Workshop files**

Current draft inputs (already in MinIO):

* `pvfs://MinIO/raw-data/csv/customers.csv`
* `pvfs://MinIO/raw-data/csv/sales.csv`
  {% endhint %}

{% hint style="info" %}
Status: **Draft**. This workshop is incomplete.
{% endhint %}

{% hint style="info" %}
Create a new transformation.

Use any of these options:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

{% tabs %}
{% tab title="1. Data Source streams" %}
{% tabs %}
{% tab title="1. Customers stream" %}
{% hint style="info" %}

#### Text file input

Use **Text file input** to read customer, sales, and event streams.
{% endhint %}

1. Drag & drop 'Text file input' steps onto the canvas.
2. Save transformation as: `customer_360.ktr` in your workshop folder.

***

**Sales (Order Management)**

1. Double-click on the first TFI step, and configure with the following properties:

| Setting          | Value                                 |
| ---------------- | ------------------------------------- |
| Step name        | `Sales`                               |
| Filename         | `pvfs://MinIO/raw-data/csv/sales.csv` |
| Delimiter        | ,                                     |
| Head row present | ✅                                     |
| Format           | mixed                                 |

<figure><img src="/files/XlXKeRhImBeHbuaSmdUZ" alt=""><figcaption><p>Select - sales.csv from VFS connections</p></figcaption></figure>

2. Click: **Get Fields** to auto-detect columns.

{% hint style="info" %}
**Business Logic:** Note that `sale_amount` may differ from `price * quantity` due to:

* Volume discounts
* Promotional pricing
* Customer-specific pricing tiers
* Currency conversion (for international sales)
  {% endhint %}

<figure><img src="/files/nLGRIikQYnpXPcuFjQy7" alt=""><figcaption><p>Get Fields - Sales</p></figcaption></figure>

3. Preview data.

<figure><img src="/files/RNrGgMfChgJ6FwVj3Qzh" alt=""><figcaption><p>Preview data - Sales</p></figcaption></figure>

{% hint style="info" %}
**Business Significance:**

* `sale_amount`: Actual revenue (may include discounts)
* `quantity`: Volume metrics for demand planning
* `payment_method`: Payment preference insights
* `status`: Filter out cancelled/refunded orders
  {% endhint %}

***

{% hint style="info" %}

#### Sort rows

{% endhint %}

1. Drag & drop 'Sort rows' steps onto the canvas.
2. Create a Hop between 'Read Customers' & 'Sort rows'.
3. Double-click on 'Sort rows' and configure the sort keys.
   {% endtab %}

{% tab title="2. Sales stream" %}
{% hint style="info" %}
Status: **Draft**. Define sales-level aggregations (for example, total spend per customer).
{% endhint %}
{% endtab %}

{% tab title="3. User Events stream" %}

{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="2. Joins" %}
{% hint style="info" %}
Status: **Draft**. Join the customer, sales, and user event streams.
{% endhint %}
{% endtab %}

{% tab title="3. Output" %}
{% hint style="info" %}
Status: **Draft**. Create one row per customer and write to `pvfs://MinIO/staging/customer360/`.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Log Parsing" %}
{% hint style="warning" %}

#### Log Parsing and Anomaly Detection

**Objective:** Parse application logs, extract metrics, and detect anomalies.

**Skills:** Regex, timestamp parsing, time-series analysis, conditional logic
{% endhint %}

<figure><img src="/files/AeSI34E4B158rfBbRU6g" alt=""><figcaption><p>Log Analysis</p></figcaption></figure>

{% hint style="info" %}
**Workshop files**

Status: **Draft**. Sample files are not published yet.
{% endhint %}

{% hint style="info" %}
Status: **Draft**. Steps coming soon.
{% endhint %}
{% endtab %}

{% tab title="Fraud" %}
{% hint style="warning" %}

#### Transactions & Fraud Detection

**Objective:** Process credit card transactions, enrich with account and merchant data, calculate transaction metrics, and detect suspicious patterns using rule-based fraud detection.

**Skills:** Financial data processing, multi-table joins, running totals, rule-based fraud detection, transaction velocity analysis

**Business Context:** A payment processor needs to analyze transaction data in real-time to detect potentially fraudulent activity before authorizing transactions. The system must flag high-risk transactions based on amount thresholds, unusual merchant activity, account balance checks, and transaction velocity patterns.
{% endhint %}

{% hint style="info" %}
**Workshop files**

Status: **Draft**. Sample files are not published yet.
{% endhint %}

{% hint style="info" %}
Status: **Draft**. Steps coming soon.
{% endhint %}
{% endtab %}

{% tab title="Data Lake Ingestion" %}
{% hint style="warning" %}

#### Data Lake Ingestion

Modern data lakes often receive the same entities (products, customers, orders) from multiple sources in different formats. This workshop demonstrates how to ingest, normalize, validate, and deduplicate multi-format data into a unified schema - a common data engineering pattern.

**Objective:** Combine data from CSV, JSON, and XML into a unified product schema.

**Skills:** Multi-format parsing, schema normalization, data validation, deduplication
{% endhint %}

{% hint style="info" %}
**Workshop files**

These files are already in MinIO:

* `pvfs://MinIO/raw-data/csv/products.csv`
* `pvfs://MinIO/raw-data/json/api_response.json`
* `pvfs://MinIO/raw-data/xml/inventory.xml`
  {% endhint %}

{% hint style="info" %}
Create a new transformation.

Use any of these options:

* Select **File** > **New** > **Transformation**
* Use `Ctrl+N` (Windows/Linux) or `Cmd+N` (macOS)
  {% endhint %}

{% tabs %}
{% tab title="1. Define Target Schema" %}
{% hint style="info" %}

#### Define Target Schema

**Objective:** Design a unified schema that accommodates all source formats.

**Why Important:** Before ingesting data, you need a clear target schema. This ensures consistency across all sources and makes downstream analytics easier.
{% endhint %}

<table data-full-width="true"><thead><tr><th width="141">Field</th><th width="109">Type</th><th width="95">Length</th><th width="125">Description</th><th>Source Mapping</th></tr></thead><tbody><tr><td>product_id</td><td>String</td><td>50</td><td>Unique product identifier</td><td>CSV: product_id<br>JSON: product_id<br>XML: sku</td></tr><tr><td>product_name</td><td>String</td><td>200</td><td>Product display name</td><td>CSV: product_name<br>JSON: product_name<br>XML: name</td></tr><tr><td>category</td><td>String</td><td>100</td><td>Product category</td><td>CSV: category<br>JSON: (derived from order type)<br>XML: category</td></tr><tr><td>price</td><td>Number</td><td>15,2</td><td>Unit price in USD</td><td>CSV: price<br>JSON: unit_price<br>XML: null (not available)</td></tr><tr><td>quantity</td><td>Integer</td><td>10</td><td>Available stock quantity</td><td>CSV: stock_quantity<br>JSON: quantity<br>XML: quantity</td></tr><tr><td>source_system</td><td>String</td><td>10</td><td>Origin system identifier</td><td>Constant: 'csv', 'json', or 'xml'</td></tr><tr><td>ingestion_time</td><td>Timestamp</td><td>-</td><td>When record was ingested</td><td>System timestamp</td></tr></tbody></table>

***

{% hint style="info" %}

#### Schema Discovery & Analysis

**Objective:** Understand each source structure before you design the target schema.

**Why it matters:** You can’t normalize what you haven’t inspected.
{% endhint %}

{% stepper %}
{% step %}
**Inspect each Data Source**

Use real samples. Avoid guessing field names.

{% tabs %}
{% tab title="CSV (products.csv)" %}
{% code title="Inspect the file" %}

```bash
mc cat minio-local/raw-data/csv/products.csv | head -5
```

{% endcode %}

{% code title="Sample output" %}

```csv
product_id,product_name,category,price,stock_quantity
PROD-001,Laptop Pro 15,Electronics,999.99,50
PROD-002,Office Chair,Furniture,299.99,100
PROD-003,Coffee Maker,Appliances,79.99,200
```

{% endcode %}

**Findings**

* Has `product_id`, `product_name`, `category`, `price`, `stock_quantity`.
* Completeness looks high.
* Naming is consistent and explicit.
  {% endtab %}

{% tab title="JSON (api\_response.json)" %}
{% code title="Inspect one nested item" %}

```bash
mc cat minio-local/raw-data/json/api_response.json | jq '.data.orders[0].items[0]'
```

{% endcode %}

{% code title="Sample output" %}

```json
{
  "product_id": "PROD-001",
  "product_name": "Laptop Pro 15",
  "unit_price": 999.99,
  "quantity": 2
}
```

{% endcode %}

**Findings**

* Has `product_id` and `product_name`.
* Uses `unit_price` instead of `price`.
* `quantity` is order quantity, not stock.
* `category` is missing.
* Path is `$.data.orders[*].items[*]`.
  {% endtab %}

{% tab title="XML (inventory.xml)" %}
{% code title="Inspect one item node" %}

```bash
mc cat minio-local/raw-data/xml/inventory.xml | grep -A 6 "<item>" | head -10
```

{% endcode %}

{% code title="Sample output" %}

```xml
<item>
  <sku>PROD-001</sku>
  <name>Laptop Pro 15</name>
  <category>Electronics</category>
  <quantity>50</quantity>
  <location>A-15</location>
</item>
```

{% endcode %}

**Findings**

* Uses `sku` for `product_id`.
* Uses `name` for `product_name`.
* Has `category` and warehouse `quantity`.
* `price` is missing.
* `location` is extra for a product master.
  {% endtab %}
  {% endtabs %}
  {% endstep %}

{% step %}
**Build a field mapping matrix**

This shows name differences and missing fields.

<table data-full-width="true"><thead><tr><th>Unified field</th><th>CSV</th><th>JSON</th><th width="109">XML</th><th>Notes</th></tr></thead><tbody><tr><td>Identifier</td><td><code>product_id</code></td><td><code>product_id</code></td><td><code>sku</code></td><td>Same meaning. Different name in XML.</td></tr><tr><td>Name</td><td><code>product_name</code></td><td><code>product_name</code></td><td><code>name</code></td><td>Same meaning. Different name in XML.</td></tr><tr><td>Category</td><td><code>category</code></td><td>❌</td><td><code>category</code></td><td>Missing in JSON.</td></tr><tr><td>Price</td><td><code>price</code></td><td><code>unit_price</code></td><td>❌</td><td>Different name in JSON. Missing in XML.</td></tr><tr><td>Stock quantity</td><td><code>stock_quantity</code></td><td><code>quantity</code></td><td><code>quantity</code></td><td>JSON <code>quantity</code> is not stock.</td></tr></tbody></table>

**What to watch**

* Missing data is normal in multi-source ingestion.
* Same name can mean different things.
  {% endstep %}

{% step %}
**Make schema decisions**

Write these down. You will forget them later.

**Field names**

* Use CSV naming as the standard.
* Map XML `sku → product_id` and `name → product_name`.
* Map JSON `unit_price → price`.

**Missing fields**

* Missing `category` in JSON: set a default like `E-commerce`.
* Missing `price` in XML: leave `NULL`.

**Data types**

* `product_id`: string. It contains `PROD-` prefix.
* `product_name`: string. Allow up to 200 chars.
* `category`: string. Allow up to 100 chars.
* `price`: decimal(15,2).
* `quantity`: integer.

**Metadata**

* Add `source_system` for lineage.
* Add `ingestion_time` for auditability.
  {% endstep %}

{% step %}
**Define a deduplication rule**

Same `product_id` can appear in multiple sources.

{% code title="Example collision" %}

```
CSV:  PROD-001, price=999.99, stock_quantity=50, category=Electronics
JSON: PROD-001, price=999.99, quantity=2,       category=NULL
XML:  PROD-001, price=NULL,   quantity=50,      category=Electronics
```

{% endcode %}

**Recommended rule**

1. Prefer CSV.
2. Then JSON.
3. Then XML.

Implement this with `source_priority` (CSV=1, JSON=2, XML=3).
{% endstep %}

{% step %}
**Checklist**

* You inspected real records for each source.
* You captured paths for nested formats.
* You documented mappings and type choices.
* You decided how to handle missing data.
* You decided how to dedupe collisions.
  {% endstep %}
  {% endstepper %}
  {% endtab %}

{% tab title="2. Ingest Data Sources" %}
{% hint style="info" %}

#### Ingest Data Sources

{% endhint %}

{% hint style="warning" %}
**Path convention used below:** `pvfs://MinIO/...`

`MinIO` is the **VFS connection name**. It must match your connection exactly.
{% endhint %}

{% stepper %}
{% step %}
**Ingest CSV products**

**Goal:** Read `products.csv` and map it to the unified schema.

**Path:** `pvfs://MinIO/raw-data/csv/products.csv`

1. Add a **Text file input** step.
   * Step name: `Read CSV Products`
   * File/directory: `pvfs://MinIO/raw-data/csv/products.csv`
   * Separator: `,`
   * Enclosure: `"` (double quote)
   * Header row present: enabled
2. On **Fields**, select **Get Fields**.
3. Add a **Select values** step.
   * Step name: `Map CSV to Target Schema`
   * Rename `stock_quantity` → `quantity`
4. Add **Add constants**.
   * Step name: `Add CSV Metadata`
   * Add field `source_system` = `csv`
5. Add **Get System Info**.
   * Step name: `Add Ingestion Timestamp`
   * Add field `ingestion_time` = `system date (variable)`

**Preview check**

* Expected rows: `12`
* `product_id`, `product_name`, `category` should be populated.
  {% endstep %}

{% step %}

### Ingest JSON order items

**Goal:** Extract product fields from nested JSON order items.

**Path:** `pvfs://MinIO/raw-data/json/api_response.json`

{% hint style="warning" %}
`quantity` in JSON is **order quantity**, not stock quantity.

Keep it as `quantity` only if that’s what you want to model.
{% endhint %}

1. Add a **JSON Input** step.
   * Step name: `Read JSON Products`
   * File: `pvfs://MinIO/raw-data/json/api_response.json`
   * Ignore empty file: enabled
2. On **Fields**, use **explicit JSONPaths** (recommended):
   * `product_id`: `$.data.orders[*].items[*].product_id`
   * `product_name`: `$.data.orders[*].items[*].product_name`
   * `unit_price`: `$.data.orders[*].items[*].unit_price`
   * `quantity`: `$.data.orders[*].items[*].quantity`

<details>

<summary>Alternative approach (base path + relative field paths)</summary>

If your PDI build supports a base “Path” for the JSON Input step, set:\n\n- Base path: `$.data.orders[*].items[*]`\n\nThen set field paths relative to the base:\n\n- `product_id`: `product_id`\n- `product_name`: `product_name`\n- `unit_price`: `unit_price`\n- `quantity`: `quantity`\n

</details>

3. Add a **Select values** step.
   * Step name: `Map JSON to Target Schema`
   * Rename `unit_price` → `price`
4. Add **Add constants**.
   * Step name: `Add JSON Metadata`
   * `source_system` = `json`
   * `category` = `E-commerce` (default)
5. Add **Get System Info**.
   * Step name: `Add JSON Ingestion Timestamp`
   * `ingestion_time` = `system date (variable)`

**Preview check**

* Expected rows: `~10–15` (can vary with sample file).
* `product_name` should not be NULL.
  {% endstep %}

{% step %}

### Ingest XML inventory items

**Goal:** Extract inventory items from XML using XPath.

**Path:** `pvfs://MinIO/raw-data/xml/inventory.xml`

1. Add **Get data from XML**.
   * Step name: `Read XML Products`
   * File: `pvfs://MinIO/raw-data/xml/inventory.xml`
   * Loop XPath: `/inventory/items/item`
2. On **Fields**, add:
   * `sku` (String)
   * `name` (String)
   * `category` (String)
   * `quantity` (Integer)

{% hint style="info" %}
Field XPaths are **relative to the loop node**.

Example: `sku` means “read the `<sku>` element under each `<item>`”.
{% endhint %}

3. Add a **Select values** step.
   * Step name: `Map XML to Target Schema`
   * Rename `sku` → `product_id`
   * Rename `name` → `product_name`
   * Add a new field `price` in **Meta-data** (type `Number`). Leave it empty (NULL).
4. Add **Add constants**.
   * Step name: `Add XML Metadata`
   * `source_system` = `xml`
5. Add **Get System Info**.
   * Step name: `Add XML Ingestion Timestamp`
   * `ingestion_time` = `system date (variable)`

**Preview check**

* Expected rows: `~8–10`
* If you get `0` rows, re-check the Loop XPath.
  {% endstep %}
  {% endstepper %}
  {% endtab %}

{% tab title="3. Merge streams" %}
{% hint style="info" %}

#### Merge Streams

**Objective:** Merge all three data streams (CSV, JSON, XML) into one unified stream.

**Why Append Streams:** This step stacks all rows from different sources vertically - like a SQL UNION ALL.
{% endhint %}

**Configuration:**

1. **Add Append streams step**
   * **Name**: "Combine All Products"
2. **Connect all three streams** to this step:
   * "Add Ingestion Timestamp" (CSV branch) → Append streams
   * "Add JSON Ingestion Timestamp" (JSON branch) → Append streams
   * "Add XML Ingestion Timestamp" (XML branch) → Append streams
3. **Important:** All input streams MUST have the same fields with the same names and types:
   * product\_id (String)
   * product\_name (String)
   * category (String)
   * price (Number) - can be null
   * quantity (Integer)
   * source\_system (String)
   * ingestion\_time (Timestamp)

**Expected Output:**

* Row count: \~30-35 rows (12 CSV + 10-15 JSON + 8-10 XML)
* All products from all sources combined
* Some products will appear multiple times (duplicates to be handled in Step 7)

**Preview Check:**

```
product_id   product_name      source_system  price
PROD-001     Laptop Pro 15     csv            999.99
PROD-002     Office Chair      csv            299.99
...
PROD-001     Laptop Pro 15     json           999.99   ← Duplicate!
PROD-005     Desk Lamp         json           45.00
...
PROD-001     Laptop Pro 15     xml            null     ← Duplicate, no price
PROD-002     Office Chair      xml            null
```

{% endtab %}

{% tab title="4. Data Validation" %}
{% hint style="info" %}

#### Data Validation

**Objective:** Validate data quality and route bad records to error handling.

**Why Important:** Multi-source data often has quality issues. Better to catch and handle them explicitly than have them cause downstream failures.
{% endhint %}

**Configuration:**

1. **Add Data Validator step**
   * **Name**: "Validate Product Data"
2. **Validations tab** - Add validation rules:

   | Fieldname     | Validation Type  | Configuration       | Error Message                   |
   | ------------- | ---------------- | ------------------- | ------------------------------- |
   | product\_id   | NOT NULL         |                     | Product ID is required          |
   | product\_id   | NOT EMPTY STRING |                     | Product ID cannot be empty      |
   | product\_name | NOT NULL         |                     | Product name is required        |
   | product\_name | NOT EMPTY STRING |                     | Product name cannot be empty    |
   | price         | NUMERIC RANGE    | Min: 0, Max: 999999 | Price must be >= 0 (if present) |
   | quantity      | NUMERIC RANGE    | Min: 0, Max: 999999 | Quantity must be >= 0           |
3. **Options tab**:
   * ☑ **Concatenate errors**: Shows all validation errors for a row
   * **Separator**: `,` (comma-space)
   * ☑ **Output all errors as one field**: `validation_errors`
4. **Add Filter rows step** after Data Validator
   * **Name**: "Route Valid vs Invalid"
5. **Condition**:

   ```
   validation_errors IS NULL
   ```

   * **True** (valid records) → Continue to deduplication
   * **False** (invalid records) → Error output
6. **Add Text file output for errors** (connect from False branch):
   * **Name**: "Write Error Records"
   * **Filename**: `pvfs://MinIO/curated/products/errors/validation_errors_${Internal.Job.Start.Date.yyyyMMdd}.csv`
   * **Include date in filename**: Helps track when errors occurred
   * **Fields to output**: All fields + `validation_errors`

**Expected Output:**

* Valid records: \~95-100% should pass (25-35 rows)
* Invalid records: 0-5% to error file (0-2 rows)

**Common Validation Failures:**

* Empty product\_id or product\_name
* Negative price or quantity values
* Non-numeric values in numeric fields
  {% endtab %}
  {% endtabs %}
  {% endtab %}
  {% endtabs %}


# SMB

File sharing ..

{% hint style="info" %}

#### **SMB/CIFS**

**Server Message Block (SMB)** is a Windows-based network file sharing protocol that enables organizations to share files, printers, and other resources across their network. Commonly used alongside the Common Internet File System (CIFS) protocol, SMB operates on a client-server model where servers provide shared file systems that clients can mount and access as if they were local disk drives. This protocol is fundamental to Windows networking environments and is supported by many enterprise storage vendors including NetApp, Dell/EMC, and Pentaho Network Attached Storage systems.

**Business Case and Benefits**

The primary business advantage of SMB connectivity is enabling secure, centralized data access across the enterprise. Organizations benefit from consolidated file storage where business-critical data resides on network file shares rather than scattered across individual workstations. This centralization improves data governance, enables consistent backup and disaster recovery policies, and facilitates collaboration by allowing multiple users to access the same files simultaneously with appropriate permissions.

SMB integration becomes strategically valuable for data integration initiatives because it allows organizations to leverage their existing Windows file infrastructure without requiring data migration or re-platforming. Companies can extract data from departmental file shares, process it through ETL pipelines, and deliver insights without disrupting established file management practices. This reduces implementation costs and accelerates time-to-value for analytics projects.

#### Pentaho Data Integration SMB Capabilities

Pentaho Data Integration provides comprehensive SMB connectivity through its Virtual File System (VFS) framework, which was enhanced in version 10.2 to support SMB files in both the Pentaho User Console and PDI client. Organizations can connect to SMB resources using either direct VFS URI addresses (formatted as `smb://<domain>;<username>:<password>@<server>:<port>/<path>`) or by creating reusable VFS connections that store connection parameters for easy access.

PDI supports SMB file ingestion across a wide range of input steps, including CSV File Input, JSON Input, Text File Output, XML Input, Parquet Input/Output, Avro Input/Output, and many others. This allows data engineers to read files directly from SMB shares, transform the data using PDI's extensive transformation capabilities, and output results to databases, cloud storage, or other SMB locations. The Get File Names and Get SubFolder names steps enable dynamic file processing, allowing transformations to scan SMB directories and process multiple files programmatically.

For enterprise data catalog initiatives, Pentaho Data Catalog can register SMB resources as data sources, automatically scan files and folders to create metadata inventories, and support data movement through Data Pipe Templates that migrate data between SMB file systems and other platforms like RDBMS, object stores, and HDFS. This integration enables organizations to maintain visibility into their SMB-based data assets while building modern data pipelines that bridge on-premises file shares with cloud and big data platforms.
{% endhint %}

<figure><img src="/files/ubqKuNxKYgrqKULJGAbV" alt="" width="308"><figcaption><p>SMB Server</p></figcaption></figure>

***


# SMB

File sharing ..

{% hint style="warning" %}
**Workshop - SMB/CIFS**

The Server Message Block (SMB) protocol is a network file sharing protocol that allows applications on a computer to read and write to files and to request services from server programs in a computer network. The SMB protocol can be used on top of its TCP/IP protocol or other network protocols

Objective of this workshop is to:

* install & configure a basic Samba server.
* share user home directories as well as provide read-write anonymous access to selected directory.
  {% endhint %}

<figure><img src="/files/ubqKuNxKYgrqKULJGAbV" alt="" width="308"><figcaption><p>SMB Server</p></figcaption></figure>

***

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="1. SMB" %}
x

x

1. Select the OS:

{% tabs %}
{% tab title="Linux" %}
x
{% endtab %}

{% tab title="Windows" %}
{% hint style="info" %}
**Test SMB Server**

Before we fire up Pentaho Data Integration, let's test:

* SMB server is up and running
* Can log into User - Bob & Alice - & Shared spaces.
  {% endhint %}

{% hint style="danger" %}
Please ensure you have completed the following setup: [SMB](/pentaho-data-integration/setup/data-sources/storage#smb)
{% endhint %}

1. Log into your Docker Desktop to check that the SMB Docker container is up and running.

<figure><img src="/files/FupGm36U35YgQQY6nUpa" alt=""><figcaption><p>Check SMB container.</p></figcaption></figure>

2. Let's test SMB server ..

```powershell
Test-NetConnection -ComputerName localhost -Port 1445  
```

<figure><img src="/files/mBIm5BJD38CqCZM5Cc20" alt=""><figcaption></figcaption></figure>

***

**SMB Shared Folders**

Let's check we have some sample data in our container/shared folder.

{% hint style="info" %}
There's a couple of ways you could do this ..!

If you have an IDE Editor installed, you can install the Docker Container Extension, see Windows 11 Pentaho Lab.
{% endhint %}

1. In the Docker Desktop UI click on the workshop-server-smb.

<figure><img src="/files/bW85eW3dTE8gf3hCQsw3" alt=""><figcaption></figcaption></figure>

2. Click on Files > scroll down to shared folder - expand to see mounted volumes.

<figure><img src="/files/XqIqAWyUEMkTjV0FnndN" alt=""><figcaption><p>Shared folders.</p></figcaption></figure>

3. Lets connect to a data source using SMB VFS in Pentaho Data Integration.
   {% endtab %}
   {% endtabs %}

x

x
{% endtab %}

{% tab title="2. Pentaho Data Integration" %}
{% hint style="info" %}
**Pentaho Data Integration**

Pentaho Data Integration utilizes Virtual File System (VFS) as the abstraction layer within the kernel to expose different filesystems.

In PDI, you can add a VFS connection and then reference that connection whenever you want to [access files or folders on your Virtual File System](https://docs.hitachivantara.com/r/xKOgM19SLuXvacAe3WhDcg/_XKq4wZRIIp4aYl7Avi5zg).
{% endhint %}

1. Select the following OS.

{% tabs %}
{% tab title="Windows" %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
**Windows - PowerShell**

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endhint %}

x

x
{% endtab %}

{% tab title="Linux" %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
**Linux**

```bash
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endhint %}

2. Create a New Transformation.
3. Drag & drop the Text file input step onto the canvas.
4. Click on the 'View' tab.
5. Highlight 'VFS Connections' and select 'New'.

<figure><img src="/files/NBmzgVGfS0HJNuy2k7x5" alt="" width="349"><figcaption></figcaption></figure>

7. Configure with the following details:

<figure><img src="/files/URe6lzr6L1kk8P2GovIw" alt=""><figcaption></figcaption></figure>

8. Click 'Test'.

<figure><img src="/files/G8sd7U7BUu3eQV6viKKw" alt="" width="308"><figcaption><p>Test connection</p></figcaption></figure>

***

{% hint style="info" %}
**Transformation - SMB File Retrieval**

Let's create a simple Transformation to onboard data via an SMB VFS connection.
{% endhint %}

1. Create the following transformation:

<figure><img src="/files/p8OfCo89ILHIHxWShILD" alt="" width="366"><figcaption><p>tr_SMB_File_Retrieval</p></figcaption></figure>

2. Double-click on Text file input > File tab
3. Click on Browse and ensure you select:

VFS Connections > SMB > Pentaho/design-tools/data-integration/samples/transformations/files/sales\_data.csv

4. Add the path.

<figure><img src="/files/63tdTv6NyYjwfL3WcWCF" alt=""><figcaption></figcaption></figure>

5. Click on Content tab & configure with the following settings:

<figure><img src="/files/7jQtl5gwyo4S0wab1Sh8" alt=""><figcaption><p>Content</p></figcaption></figure>

6. Click on Fields tab & click on 'Get Fields'

<figure><img src="/files/68Ef1X1g7klBPZklLt3p" alt=""><figcaption><p>Get Fields</p></figcaption></figure>

7. Preview the rows.

<figure><img src="/files/RzT0rpCRd90EYZLncTNX" alt=""><figcaption><p>Preview rows</p></figcaption></figure>

8. Click OK.

{% hint style="info" %}
Add the other steps to format / rename some fields, before output as a .txt in the same directory as your Transformation.
{% endhint %}

x
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}

x

{% tabs %}
{% tab title="First Tab" %}
x
{% endtab %}

{% tab title="Second Tab" %}
x
{% endtab %}
{% endtabs %}

x


# Hitachi Content Platform

S3 Compatible Object Storage ..

{% hint style="danger" %}
This workshop is currently under development.
{% endhint %}

Hitachi Content Platform (HCP) is an object-storage solution designed for efficient and secure data management. It allows organizations to store, protect, and retrieve vast amounts of unstructured data with ease.

HCP integrates seamlessly with various applications and provides advanced features such as data deduplication, compression, and encryption. Its scalable architecture and robust governance capabilities make it suitable for both on-premises and cloud environments, ensuring data integrity and accessibility.

<figure><img src="/files/Kqb68QiEz67tYY9P4A9e" alt=""><figcaption><p>HCP Solution</p></figcaption></figure>

HCP stores objects in a repository. Each object permanently associates data HCP receives (for example, a document, an image, or a movie) with information about that data called metadata.

<figure><img src="/files/os9c1DDXcUBrrjLLt27E" alt=""><figcaption><p>Metadata</p></figcaption></figure>

In PDI, you can query the metadata to locate and access HCP objects. The HCP object consists of a read-only file, a unique URL, system metadata properties, and custom metadata annotations.

{% tabs %}
{% tab title="1. Create a VFS connection to HCP" %}
{% hint style="info" %}
A VFS (Virtual File System) connection allows you to integrate and manage different storage systems within PDI, abstracting the complexities of underlying protocols. It provides a unified interface to access a variety of storage backends like Amazon S3, Azure Data Lake, Google Cloud Storage, and more.
{% endhint %}

**Create a VFS connection**

Perform the following steps to create a VFS connection in PDI:

1. Start the PDI client (Spoon) and create a new transformation or job.
2. In the View tab of the Explorer pane, right-click on the VFS Connections folder, and then click New. The New VFS connection dialog box opens.
3. In the Connection name field, enter a name that uniquely describes this connection. The name can contain spaces, but it cannot include special characters, such as #, $, and %.
4. In the Connection type field, select from one of the following types:Amazon S3 / MinIO:/HCP (Default):
   * Simple Storage Service (S3) accesses the resources on Amazon Web Services.
   * MinIO accesses data objects on an Amazon compatible storage server.
   * HCP uses the S3 protocol to access HCP. See [Access to HCP REST](https://docs.hitachivantara.com/r/W5Oy5NghPggWbWs_8gNJXw/EI9PXAhtM9hivQPVrxVXEA) for more information.

x
{% endtab %}

{% tab title="2. Query Metadata" %}
x

x

x

x

x

x
{% endtab %}
{% endtabs %}

x

x


# Big Data

{% hint style="info" %}
Big Data refers to extremely large and complex datasets that traditional data processing tools can't handle effectively. It's characterized by the "six Vs":

Think of social media posts, sensor readings, transaction records, and video files all being created simultaneously across millions of devices.

The challenge with Big Data isn't just storing these massive datasets, but extracting meaningful insights from them quickly enough to be useful. Organizations use specialized technologies like distributed computing systems (such as Hadoop and Spark) and cloud platforms to process and analyze this information. Machine learning algorithms help identify patterns that would be impossible for humans to spot manually.

Big Data has transformed how businesses operate and make decisions. Companies use it for everything from predicting customer behavior and optimizing supply chains to detecting fraud and personalizing recommendations. In healthcare, it helps analyze patient records and research data to improve treatments. The key value lies not in having lots of data, but in using advanced analytics to turn that data into actionable intelligence that drives better outcomes.
{% endhint %}

<figure><img src="/files/9OhCXPEL0LMOzyK0iE0j" alt=""><figcaption></figcaption></figure>

***

{% hint style="info" %}
**Workshops**
{% endhint %}

{% tabs %}
{% tab title="Apache Hadoop" %}
x

x

{% content-ref url="/pages/MeTNsXZW0WrnyqieOUKE" %}
[Apache Hadoop](/pentaho-data-integration/data-integration/data-sources/big-data/apache-hadoop)
{% endcontent-ref %}
{% endtab %}

{% tab title="Snowflake" %}

{% endtab %}
{% endtabs %}

x


# Snowflake

Snowflake is an analytic data warehouse running completely on a cloud infrastructure. Snowflake supports loading popular data formats like JSON, Avro, Parquet, ORC, and XML. Using Pentaho Data Integration (PDI), you can load your data into Snowflake and define jobs in PDI to efficiently orchestrate warehouse operations, paying only for the storage and computing resources actually used when you use them.

Using the Snowflake job entries in PDI, data engineers can set up virtual warehouses, bulk load data, and stop the warehouse when the process is complete. You can scale Snowflake virtual warehouses up and down, or suspend them when not in use to reduce costs.

In this Lab you'll be using Pentaho Data Integration with Snowflake to:

* leverage PDI's broad set of transformation capabilities to ingest and prepare data before loading it into Snowflake for advanced analytics.

<figure><img src="/files/8uoprjSga6YALW5c2stY" alt=""><figcaption><p>Snowflake -AWS</p></figcaption></figure>


# Apache Hadoop

Big Data stuff ..

{% hint style="info" %}
**Hadoop**

Hadoop is a highly scalable, open-source, distributed computing platform that allows you to store and process large amounts of data. It is used to store and analyze data from a variety of sources, including databases, web servers, and file systems.

Hadoop is designed to be scalable by distributing the processing of data across a large number of computers. It also allows you to store and analyze data in a way that is faster than traditional methods. Hadoop is used to store and analyze data from a variety of sources, including databases, web servers, and file systems.
{% endhint %}

<figure><img src="/files/1RLI8MgJF1vf2Onr8cuY" alt=""><figcaption><p>Apache Hadoop Architecture</p></figcaption></figure>

{% embed url="<https://hadoop.apache.org/>" %}
Link to Apache Hadoop
{% endembed %}

#### Components

{% tabs %}
{% tab title="Master / Slave / Data Nodes" %}
{% hint style="info" %}
The master node keeps track of the status of all the data nodes. If a data node goes down, the master node takes over the processing of that block. The slave nodes process the data on their own. HDFS requires a high-speed Internet connection. It is usually best to have at least a 10 Mbps network connection.

HDFS works on a time-based algorithm, which means that every block is processed in a predetermined time interval. This provides a high degree of scalability as all the nodes process the data in real time. HDFS is a great solution for data warehouse and business intelligence tasks. It is also a good solution for high-volume, high-frequency data analysis.
{% endhint %}

<figure><img src="/files/MPTUG4bumZjuVvT18g6s" alt=""><figcaption><p>Master / Slave Nodes</p></figcaption></figure>

{% hint style="info" %}
When a client connection is received from a master server, the NameNode running on the master server looks into both the local namespace and the master server’s namespace to find the matching records. The NameNode running on the master server then executes the lookup and returns a list of records that match the query.

DataNodes then get the records and start storing them. A data block is a record of data. Data nodes use the data blocks to store different types of data, such as text, images, videos, etc. NameNode maintains data node connections with clients based on the replication status of DataNodes. If a DataNode goes down, the client can still communicate with the NameNode. The client then gets the latest list of data blocks from the NameNode and communicates with the DataNode that has the newest data block.

Thus, DataNode is a compute-intensive task. It is therefore recommended to use the smallest possible DataNode. As DataNode stores data, it is recommended to choose a node that is close to the centre of the data. In a distributed system, all the nodes have to run the same version of Java, so it is recommended to use open-source Java. If there are multiple DataNodes, then they are expected to work in tandem to achieve a certain level of performance.
{% endhint %}
{% endtab %}

{% tab title="Block in HDFS" %}
{% hint style="info" %}
A file stored in Hadoop has the default 128 MB or 256 MB block size.
{% endhint %}

<figure><img src="/files/Kvq4TG2exSiMzkYbkS4z" alt=""><figcaption><p>Block size</p></figcaption></figure>

{% hint style="info" %}
We have to ensure that the storage consumed by our application doesn’t exceed the level of data storage. If the storage consumed by our application is too much, then we have to choose our block size to avoid excessive metadata growth. If you are using a standard block size of 16KB, then you are good to go. You won’t even have to think about it.

Our HDFS block size will be chosen automatically. However, if we are using a large block size, then we have to think about the metadata explosion. We can either keep the metadata as small as possible or keep it as large as possible. HDFS has the option of keeping the metadata as large as possible.
{% endhint %}

***

**Replication Management**

{% hint style="info" %}
When a node fails, the data stored on it is copied to another healthy node. This process is known as “replication”. If a DataNode fails, the data stored on it is not lost. It is simply stored on another DataNode. This is a good thing because it helps in the high availability of data. When you are using a DataNode as the primary storage for your data, you must keep in mind that it is a highly-available resource. If the primary storage for your data is a DataNode, you must make sure to have a proper backup and restoration mechanism in place.

A given file can have a replication factor of 3 or 1, but it will require 3 times the storage if we keep using a replication factor of 3. The NameNode keeps track of each data node’s block report and whenever a block is under-or over-replicated, it adds or removes replicas accordingly.
{% endhint %}

<figure><img src="/files/3fdo2F6JOriDQsT6r0aV" alt=""><figcaption><p>Block Replication</p></figcaption></figure>

{% hint style="info" %}
A given file can have a replication factor of 3 or 1, but it will require 3 times the storage if we keep using a replication factor of 3. The NameNode keeps track of each data node’s block report and whenever a block is under-or over-replicated, it adds or removes replicas accordingly.
{% endhint %}

***

**Rack Awareness**

{% hint style="info" %}
When a block is deleted from a rack, the next available block will be placed on the rack. When a block is updated, the previous block is automatically updated in the same rack. This ensures data consistency across the storage network. In case of a fault in any of the data storage nodes, other storage nodes can be alerted through the rack awareness algorithm to take over the responsibility of the failed node.

This helps in providing failover capability across the storage network. This helps in providing high availability of data storage. As data stored in HDFS is massive, it makes sense to use multiple storage nodes for high-speed data access.

HDFS uses a distributed data store architecture to provide high-speed data access to its users. This distributed data store architecture allows for parallel processing of data in multiple data nodes. This parallel processing of data allows for high-speed storage of large data sets.
{% endhint %}

<figure><img src="/files/2x5wUIoOrGFvvytgafrl" alt=""><figcaption><p>Rack Awareness</p></figcaption></figure>
{% endtab %}

{% tab title="MapReduce" %}
{% hint style="info" %}
MapReduce is a data processing language and software framework that allows you to process large amounts of data in a reliable and efficient manner. It is a great fit for data-intensive, real-time and/or streaming applications. Basically, MapReduce allows you to partition your data and process items only when they are needed.

A map-reduce job consists of several maps and reduces functions. Each map function generates, parses, transforms, and filters data before passing it on to the next function. The reduced function groups, aggregates, and partitioning of this intermediate data from the map functions. The map task runs on the same node as the input source. Map tasks are responsible for generating summary statistics about data in the form of a report. The report can be viewed on a web browser or printed.
{% endhint %}

<figure><img src="/files/Boe0fLC9YKcRlDZW9XV5" alt=""><figcaption><p>MapReduce</p></figcaption></figure>

{% hint style="info" %}
The output of a map task is the same as that of a reduced task. The only difference is that the map task returns a result whereas the reduced task returns a data structure that is applicable for further analysis. A map task is usually repetitive and is triggered when the data volume on the source is greater than the volume of data that can be processed in a short period of time.
{% endhint %}
{% endtab %}

{% tab title="YARN" %}
{% hint style="info" %}
YARN, also known as Yet Another Resource Negotiator, is a resource management and job scheduling/monitoring daemon in Hadoop. A resource manager in YARN isolates resource management from job scheduling and monitoring. A global Resource Manager oversees operations for the entire YARN network, including per-application Application Master. A job or a DAG of jobs may be defined as an application.

The Resource Manager manages resources for all the competing applications in the YARN framework. The Node Manager monitors resource usage by the container and passes it on to Resource Manger. There are resources such as CPU, memory, disk, and connectivity, among others. To perform and monitor the application, the Applcation Master talks to the ResourceManager and the Node Manager to handle and manage resources.
{% endhint %}
{% endtab %}
{% endtabs %}

***


# Apache Hadoop

{% hint style="warning" %}

#### Workshop - Apache Hadoop

{% endhint %}

x

x

{% tabs %}
{% tab title="First Tab" %}
{% hint style="info" %}
**Start Container**

Hopefully .. you've completed the [Setup: Apache Hadoop](/pentaho-data-integration/setup/data-sources/big-data/apache-hadoop)
{% endhint %}

1. To start the container enter the following:

```docker
docker start -ai AHW
```

Once the Docker shell opens, just type `restart` to restart all processes.

<figure><img src="/files/0n9qW3xFOmtx74hIKvwu" alt=""><figcaption></figcaption></figure>

{% tabs %}
{% tab title="NameNode" %}
{% hint style="info" %}
**NameNode**

The **NameNode** is the master node and central component of Hadoop's Distributed File System (HDFS). It acts as the "brain" of the file system.
{% endhint %}

1. Log into NameNode:

{% embed url="<http://localhost:9870>" %}

2. You can upload files to the root directory:

<figure><img src="/files/5Nw9qB5XgCJRIRfg5UCo" alt=""><figcaption><p>Browse file system</p></figcaption></figure>
{% endtab %}

{% tab title="YARN" %}
{% hint style="info" %}
**YARN**

YARN acts as the operating system for Hadoop clusters by separating resource management from job scheduling and monitoring, allowing multiple data processing engines like MapReduce, Spark, Hive, and others to run simultaneously on the same cluster.
{% endhint %}

1. Log into YARN:

{% embed url="<http://localhost:8095>" %}

<figure><img src="/files/L9yFjdkUQXhsbRYVJscF" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="DataNode" %}
{% hint style="info" %}
**DataNode**

A DataNode in Hadoop is a worker node in the Hadoop Distributed File System (HDFS) that stores the actual data blocks and serves read/write requests from clients. DataNodes communicate regularly with the NameNode through heartbeat messages to report their health status and the blocks they're storing.

They handle data replication by creating multiple copies of blocks across different nodes to ensure fault tolerance, and they perform block verification to detect corruption. DataNodes also participate in data pipeline operations during file writes and coordinate with other DataNodes to maintain data integrity and availability across the distributed cluster.
{% endhint %}

1. Log into the DataNode:

{% embed url="<http://localhost:9864>" %}

2. Useful for troubleshooting the Node.

<figure><img src="/files/INs6WiszLQoE6jjbLb3Q" alt=""><figcaption><p>DataNode</p></figcaption></figure>

x
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

x


# Jupyter Notebook

Jupyter notebooks are used for data science tasks such as exploratory data analysis (EDA), data cleaning and transformation, data visualization, statistical modeling, machine learning, and so on ..

{% hint style="info" %}

#### **Jupyter Notebook**

Jupyter Notebook's cell-based interface creates an ideal environment for consuming and analyzing data processed through Pentaho Data Integration (PDI). The interactive coding structure allows data scientists to immediately visualize and explore PDI outputs, with results appearing below each executed cell, while documenting their analytical process through integrated Markdown explanations. This makes Jupyter the perfect downstream tool for leveraging PDI's data preparation work, enabling seamless transition from engineered datasets to advanced analytics and model development.
{% endhint %}

<figure><img src="/files/3mhqfMbQaO6GAlCyqRdR" alt=""><figcaption><p>Jupyter Notebook</p></figcaption></figure>

{% hint style="info" %}
The integration between PDI and Jupyter Notebook represents a powerful approach to enterprise data science that maximizes organizational efficiency. PDI serves as the robust data preparation engine, handling complex data blending, cleansing, and feature engineering operations that can be easily scaled and deployed to production environments.

These prepared datasets then flow seamlessly into Jupyter Notebook environments where data scientists can focus on their core expertise: model exploration, hyperparameter tuning, and advanced machine learning techniques. The notebook format perfectly complements PDI's structured outputs by providing an interactive workspace for hypothesis testing, visualization, and iterative model development.

This PDI-to-Jupyter workflow creates substantial competitive advantages for organizations. The clear separation of concerns accelerates time-to-market by allowing data engineers to optimize data pipelines in PDI while data scientists simultaneously develop models in Jupyter using previously processed datasets.

Solution quality improves through specialized tool usage, and team collaboration is enhanced as PDI's standardized outputs can be easily shared and consumed across multiple Jupyter environments. Most importantly, this integration reduces the data preparation burden on data scientists, allowing them to dedicate more time to advanced analytics while ensuring that data engineering work is properly leveraged throughout the organization's analytical workflows.
{% endhint %}


# PDI to Jupyter Notebook

{% hint style="warning" %}

#### **Workshop - PDI to Jupyter Notebook**

This workshop demonstrates how to create a Pentaho Data Integration (PDI) pipeline that processes sales data and automatically triggers analysis in Jupyter Notebook when the output file is saved.

The topics were going to cover:

* Creating a Jupyter Notebook
* Installing required Python packages: `jupyter`, `watchdog`, `xslxwriter`
* Create a PDI pipeline: sales\_data.csv file
* Create a File Watcher script
  {% endhint %}

<figure><img src="/files/Et4PZ08PY5A3CQP8MzUU" alt=""><figcaption><p>Pipeline</p></figcaption></figure>

{% hint style="info" %}
Quick overview of the pipeline:

* Execute a PDI pipeline with sample sales\_data.csv - from datasets folder
* The file output to the pdi-output folder triggers the Jupyter Notebook to
* Load the data - csv files from pdi-output - analyze and visualize the results
* Export the results to the reports folder
  {% endhint %}

***

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

Select the Host Docker OS:

{% tabs %}
{% tab title="Linux" %}
{% hint style="info" %}

#### PDI to Jupyter Notebook

{% endhint %}

{% tabs %}
{% tab title="1. Setup Verification" %}
{% hint style="info" %}

#### Setup Verification

Before building the PDI pipeline, verify everything works by running the sample notebook.
{% endhint %}

1. Verify Python Packages are Installed.

```bash
# Check that packages were auto-installed
docker exec jupyter-datascience pip list | grep -E "watchdog|xlsxwriter"
# Expected: watchdog x.x.x  and  XlsxWriter x.x.x
```

{% hint style="info" %}
Python packages (`watchdog`, `xlsxwriter`) are **automatically installed** when the container starts via the `post-start.sh` startup script.

If packages are missing, check the container logs: `docker logs jupyter-datascienc`
{% endhint %}

2. Verify Test Files Exist (Inside the Container).

```bash
# Still inside the container shell:
docker exec -it jupyter-datascience sh

# Check for datasets
cd /home/jovyan/datasets
ls
# Expected: orders.csv  sales_data.csv

# Check for notebooks
cd /home/jovyan/notebooks
ls
# Expected: sales_analysis.ipynb  welcome.ipynb

# Exit the container shell
exit
```

***

**Run the Sales Analysis Notebook**

1. In Jupyter Lab, navigate to **notebooks/** in the file browser
2. Open **sales\_analysis.ipynb**
3. Run each cell in order (Shift+Enter or use the Run menu)
4. The notebook will:
   * Load `sales_data.csv` from `/home/jovyan/datasets/`
   * Generate a 4-panel Sales Analysis Dashboard
   * Calculate Key Metrics (revenue, average order value, profit margin)
   * Export an Excel report to `/home/jovyan/reports/`

<figure><img src="/files/wNRRCM2gOzUig6PfPMKy" alt=""><figcaption><p>sales_analysis.ipynb</p></figcaption></figure>

5. Check the Output Report

```bash
# On the host machine, check for the generated report
ls ~/Jupyter-Notebook/reports/
# Expected: sales_analysis_<timestamp>.xlsx
```

Open the Excel file and verify it has two sheets:

* **Summary** - Key metrics (Total Revenue, Average Order Value, etc.)
* **Detailed Data** - Full processed dataset

<figure><img src="/files/4JBspnYq14DVoZl1aDC7" alt=""><figcaption><p>sales_analysis</p></figcaption></figure>
{% endtab %}

{% tab title="2. PDI pipeline" %}
{% hint style="info" %}

#### Build the PDI Pipeline

The data scientists have deployed the sales\_analysis.ipynb notebook. The notebook will be triggered by a File Watcher that's polling the \~/Jupyter-Notebook/pdi-output/ for:

* sales\_detailed\_\*.csv

So in this part of the workshop, we're going to create a simple pipeline that:

* Loads the sales.csv
* Cleans and performs some calculations and aggregations
* Outputs to: \~/Jupyter-Notebook/pdi-output/ folder.
  {% endhint %}

1. Start Pentaho Data Integration (Spoon)

```bash
# Navigate to the PDI installation
cd
cd ~/Pentaho/design-tools/data-integration

# Launch Spoon (the PDI graphical designer)
./spoon.sh
```

2. Create a New Transformation - sales\_pipeline.ktr:

<figure><img src="/files/wc4yXOVvhmStK31dnZJa" alt=""><figcaption><p>sales_pipeline.ktr</p></figcaption></figure>

{% file src="/files/WmBEitG9qAyyARlGxF7q" %}

{% file src="/files/0MzlYl8FsUkRYq0f4K9F" %}

{% tabs %}
{% tab title="1. CSV File input" %}
{% hint style="info" %}

#### **CSV File input**

The CSV File Input transform extracts data from delimited files using either a predefined schema or manually configured field layouts. Despite its name, this transform supports any delimiter—pipes, tabs, semicolons, or custom separators—not just commas.

Built for speed through optimized internal processing, this transform offers a focused subset of Text File Input capabilities with three key performance advantages:

**Native I/O (NIO)** uses direct system calls for faster file reading, though it's currently limited to local files without VFS support.

**Parallel Processing** enables distributed file reading when running multiple transform copies or in clustered mode. Each copy processes a separate file block, allowing workload distribution across multiple threads or slave nodes.

**Lazy Conversion** optimizes performance for pass-through data scenarios. When fields flow unchanged from input to output (like file-to-database transfers), this feature prevents unnecessary data type conversions, avoiding the overhead of converting raw data into strings, dates, or numbers.

While this transform has fewer configuration options than the general Text File Input transform, these performance optimizations make it ideal for high-throughput data processing workflows.
{% endhint %}

1. Drag & drop a CSV File input step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/HZkBxXDkqUnEG6PGqOVs" alt=""><figcaption><p>CSV file input</p></figcaption></figure>

**CSV File Input**

* **Step type:** Input > CSV file input
* **Purpose:** Reads the source sales data
* **Configuration:**
  1. Drag a **CSV file input** step onto the canvas
  2. Double-click to configure:
     * **Filename:** `~/Jupyter-Notebook/datasets/sales_data.csv`
     * **Delimiter:** `,`
     * **Header row present:** checked
  3. Click **Get Fields** to auto-detect the 8 columns
  4. Verify the field types: `order_id` (Integer), `customer_id` (Integer), `product_name` (String), `product_category` (String), `quantity` (Integer), `unit_price` (Number), `cost` (Number), `order_date` (String)
  5. Click **Preview** to verify data loads correctly (should show 250 rows)
     {% endtab %}

{% tab title="2. Data Validator" %}
{% hint style="info" %}

#### **Data Validator**

The Data Validator transform enables you to define validation rules that check input data across different fields in each row. When the validator encounters a row that violates one or more validation rules, it generates an error or exception.

You can capture all validation errors by configuring an error handling hop from this transform, which will provide you with a comprehensive list of any validation failures that occur during processing.
{% endhint %}

1. Drag & drop Data Validator step onto the canvas.
2. Double-click on the step, and configure the following properties:

**Validation: quantity**

<figure><img src="/files/3KXKZklsMiwMGtZjtxZm" alt=""><figcaption><p>Validation - quantity</p></figcaption></figure>

**Validation: unit\_price**

<figure><img src="/files/XQzzpBYuf0UgkW0Ft1Zm" alt=""><figcaption><p>Validation - unit_price</p></figcaption></figure>

**Data Validator (optional)**

* **Step type:** Validation > Data Validator
* **Purpose:** Validates input data quality
* **Configuration:**
  1. Add a **Data Validator** step and connect it from CSV file input
  2. Add validations:
     * `quantity`: Data type = Integer, Minimum value = 1
     * `unit_price`: Data type = BigNumber, Minimum value = 0.01
  3. Create an **error handling hop** (red hop) to a **Dummy** step for invalid records
     {% endtab %}

{% tab title="3. Replace in String" %}
{% hint style="info" %}

#### **Replace in String**

The Replace In String transform performs simple search and replace operations on String fields.

It also supports regular expressions and group references. group references are picked up in the replace by string as $n where n is the number of the group.

You can also use hexadecimal replacements for special characters.
{% endhint %}

1. Drag & drop Replace in string step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/abqbGZcT7JLWJ72FrIHt" alt=""><figcaption><p>Replace in string - standardize</p></figcaption></figure>

**Replace in String**

* **Step type:** Transform > Replace in String
* **Purpose:** Normalise the inconsistent `product_category` casing
* **Configuration:**
  1. Add a **Replace in String** step
  2. Add rows to fix each category (the dataset has mixed case like `electronics` / `Electronics`):

     <table><thead><tr><th width="174">In stream field</th><th width="122">Search</th><th width="133">Replace with</th><th width="120">Use RegEx</th></tr></thead><tbody><tr><td>product_category</td><td><code>electronics</code></td><td><code>Electronics</code></td><td>no</td></tr><tr><td>product_category</td><td><code>clothing</code></td><td><code>Clothing</code></td><td>no</td></tr><tr><td>product_category</td><td><code>home</code></td><td><code>Home</code></td><td>no</td></tr><tr><td>product_category</td><td><code>sports</code></td><td><code>Sports</code></td><td>no</td></tr><tr><td>product_category</td><td><code>books</code></td><td><code>Books</code></td><td>no</td></tr></tbody></table>
  3. Set **Case sensitive** to `no` for each row to catch all variations
     {% endtab %}

{% tab title="4. Calculator" %}
{% hint style="info" %}

#### **Calculator**

The Calculator transform offers built-in functions that operate on your input field values.

Beyond the function arguments (Field A, Field B, and Field C), you need to define the return type for each function.

After calculations are complete, you can exclude specific fields from the output using the Remove option. This feature helps eliminate temporary values that aren't needed in your final pipeline.

The Calculator delivers significantly faster execution compared to custom JavaScript scripts.
{% endhint %}

1. Drag & drop calculator step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/AGHLdqvc6qoKA35R516L" alt=""><figcaption><p>Calculator</p></figcaption></figure>

**Calculator**

* **Step type:** Transform > Calculator
* **Purpose:** Compute derived fields
* **Configuration:**
  1. Add a **Calculator** step
  2. Add two calculations:

     | New field      | Calculation | Field A    | Field B      |
     | -------------- | ----------- | ---------- | ------------ |
     | `total_amount` | A \* B      | `quantity` | `unit_price` |
     | `total_cost`   | A \* B      | `quantity` | `cost`       |
  3. To compute profit margin, add a **User Defined Java Expression** step (or a second Calculator step) after this one:
     * `profit_margin` = `(total_amount - total_cost) / total_amount`
       {% endtab %}

{% tab title="5. Formula" %}
{% hint style="info" %}

#### **Formula**

The Formula step can calculate Formula Expressions within a data stream. It can be used to create simple calculations like \[A]+\[B] or more complex business logic with a lot of nested if / then logic.
{% endhint %}

1. Drag & drop Formula step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/9dejD4TaivkcgrjYg0GM" alt=""><figcaption><p>Formula - profit margin</p></figcaption></figure>

**Formula**

* **Step type:** Scripting > Formula
* **Purpose:** Calculate profit margin using Pentaho Formula Engine (Libformula)
* **Configuration:**
  1. Add a **Formula** step and connect it from the Calculator step
  2. Click **Add** to create a new formula:

     | Field name      | Formula                                          | Value type |
     | --------------- | ------------------------------------------------ | ---------- |
     | `profit_margin` | `[total_amount] - [total_cost] / [total_amount]` | Number     |
  3. The formula uses field references in square brackets (e.g., `[total_amount]`)
  4. This calculates the profit margin as a decimal (e.g., 0.35 = 35% margin)

{% hint style="info" %}
The Formula step uses Pentaho's Libformula engine which supports Excel-like formulas. For simple arithmetic like this, you could also use a Calculator step or User Defined Java Expression. However, Formula steps are more flexible for complex calculations.
{% endhint %}
{% endtab %}

{% tab title="6. Text file output" %}
{% hint style="info" %}

#### **Text file output**

The Text File Output transform exports data to text file formats, most commonly generating CSV files that can be opened in spreadsheet applications like Excel.

This transform also supports creating fixed-width files by specifying field lengths in the fields configuration tab. You have two options for defining the output structure: use an existing Schema Definition or manually configure the field layout. When working with a Schema Definition, pair this transform with the Schema Mapping transform to align your incoming data stream with the chosen schema structure.
{% endhint %}

1. Drag & drop Text file output step onto the canvas.
2. Double-click on the step, and configure the following properties:

<figure><img src="/files/kCugISBUUKErwcVY9zzf" alt=""><figcaption><p>Text file output</p></figcaption></figure>

**Text File Output**

* **Step type:** Output > Text file output
* **Purpose:** Write the processed data to the pdi-output folder
* **Configuration:**
  1. Add a **Text file output** step
  2. Configure the File tab:
     * **Filename:** `~/Jupyter-Notebook/pdi-output/sales_detailed`
     * **Extension:** `csv`
     * **Include date in filename:** Yes
     * **Date time format:** `yyyyMMdd_HHmmss` (produces `sales_detailed_20250218_143022.csv`)
  3. Configure the Content tab:
     * **Separator:** `,`
     * **Header:** Yes
  4. Click **Get Fields** to populate the output field list
     {% endtab %}
     {% endtabs %}
     {% endtab %}

{% tab title="3. File watcher" %}
{% hint style="info" %}

#### File watcher

The file watcher monitors the `pdi-output/` directory and **automatically executes** the analysis notebook inside the Docker container when PDI writes a new file.
{% endhint %}

1. Start File watcher - in a new terminal.

```bash
# Navigate to the scripts directory
cd
cd ~/Jupyter-Notebook/scripts/

# Create a Python virtual environment and install watchdog
# (Modern Linux distros block system-wide pip installs - PEP 668)
python3 -m venv .venv
.venv/bin/pip install watchdog

# Start the file watcher using the venv Python
.venv/bin/python3 file_watcher.py
```

**Expected output:**

```
Watching folder: /home/<user>/Jupyter-Notebook/pdi-output
Press Ctrl+C to stop...
```

2. Re-run the transformation. The file watcher detects the new `sales_detailed.csv` and **auto-executes** the notebook:

<figure><img src="/files/wu3SeUmKy5JHObTInKSc" alt=""><figcaption><p>File watcher</p></figcaption></figure>

{% hint style="info" %}
The file watcher uses `docker exec` to run `jupyter nbconvert --execute` inside the container, so the notebook runs automatically without you having to open Jupyter Lab.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}

{% tab title="Windows" %}
x

x

x

x

{% tabs %}
{% tab title="1. Jupyter Notebook" %}
{% hint style="info" %}
**Quick Setup**

To check the various scripts and that volume mappings are working, let's analyze a sample sales\_data.csv:

* Install some python packages
* Load a sample dataset - test\_sales\_data.csv
* Run the sales\_analysis.ipynb - check container paths
* Check ouput
  {% endhint %}

{% hint style="danger" %}
Please ensure you have completed the following setup: [Jupyter Notebook](/pentaho-data-integration/setup/data-sources/jupyter-notebook).

Check Jupyter Notebook is running in a Docker container ..!
{% endhint %}

{% hint style="info" %}
To list / install python packages:

```docker
cd \
docker exec -it jupyter-datascience bin/bash
```

Once inside the container:

```
pip list - will list the installed packages
```

{% endhint %}

***

1. Install required Python packages:

```
cd \
docker exec -it jupyter-datascience bash
pip install jupyter watchdog xlsxwriter
```

2. Check for the test\_sales\_data.csv & sales\_analysis.ipynb (still in container):

```bash
cd
cd /home/jovyan/datasets
ls
```

```bash
cd
cd /home/jovyan/notebooks
ls
```

3. Open the sales\_analysis.ipynb notebook and RUN each section:

<figure><img src="/files/DbwjvCEhU755peQWs14x" alt=""><figcaption><p>RUN the Notebook</p></figcaption></figure>

4. Check for reports: C:\Jupyter-Notebook\reports\sales\_analysis\_timestamp.xlsx

<figure><img src="/files/4uoibygQV5u5MbqJrB4G" alt=""><figcaption><p>Reports</p></figcaption></figure>

{% hint style="warning" %}
Check you have 2 sheets: Summary & Detailed Data.
{% endhint %}
{% endtab %}

{% tab title="2. File Watcher" %}
{% hint style="info" %}
**File Watcher**
{% endhint %}

x

x

x
{% endtab %}

{% tab title="3. Pentaho Data Integration" %}
{% hint style="info" %}
**Data Pipeline**

The data scientists have deployed the sales\_analysis.ipynb notebook. The notebook will be triggered by a File Watcher that's polling the C:\Jupyter-Notebook\pdi-output for:

* sales\_detailed\_\*.csv

So in this part of the workshop, we're going to create a simple pipeline that:

* Loads the sales.csv
* Cleans and performs some calculations and aggregations
* Outputs to: C:\Jupyter-Notebook\pdi-output folder.
  {% endhint %}

1. Start Pentaho Data Integration.

{% hint style="info" %}
**Windows - PowerShell**

```powershell
Set-Location C:\Pentaho\design-tools\data-integration
.\spoon.bat
```

{% endhint %}

{% hint style="info" %}
**Linux**

```bash
cd
cd ~/Pentaho/design-tools/data-integration
./spoon.sh
```

{% endhint %}

2. Create a New Transformation:

{% tabs %}
{% tab title="CSV File input" %}
{% hint style="info" %}
**CSV File input**

The CSV File Input transform extracts data from delimited files using either a predefined schema or manually configured field layouts. Despite its name, this transform supports any delimiter—pipes, tabs, semicolons, or custom separators—not just commas.

Built for speed through optimized internal processing, this transform offers a focused subset of Text File Input capabilities with three key performance advantages:

**Native I/O (NIO)** uses direct system calls for faster file reading, though it's currently limited to local files without VFS support.

**Parallel Processing** enables distributed file reading when running multiple transform copies or in clustered mode. Each copy processes a separate file block, allowing workload distribution across multiple threads or slave nodes.

**Lazy Conversion** optimizes performance for pass-through data scenarios. When fields flow unchanged from input to output (like file-to-database transfers), this feature prevents unnecessary data type conversions, avoiding the overhead of converting raw data into strings, dates, or numbers.

While this transform has fewer configuration options than the general Text File Input transform, these performance optimizations make it ideal for high-throughput data processing workflows.
{% endhint %}

1. Drag & drop a CSV File input step onto the canvas.
2. Double-click on the step, and configure the following properties:

x

x
{% endtab %}

{% tab title="Data Validator" %}
{% hint style="info" %}
**Data Validator**

The Data Validator transform enables you to define validation rules that check input data across different fields in each row. When the validator encounters a row that violates one or more validation rules, it generates an error or exception.

You can capture all validation errors by configuring an error handling hop from this transform, which will provide you with a comprehensive list of any validation failures that occur during processing.
{% endhint %}

x

x

x
{% endtab %}

{% tab title="Replace in String" %}
{% hint style="info" %}
**Replace in String**

The Replace In String transform performs simple search and replace operations on String fields.

It also supports regular expressions and group references. group references are picked up in the replace by string as $n where n is the number of the group.

You can also use hexadecimal replacements for special characters.
{% endhint %}

x

x

x
{% endtab %}

{% tab title="Calculator" %}
{% hint style="info" %}
**Calculator**

The Calculator transform offers built-in functions that operate on your input field values.

Beyond the function arguments (Field A, Field B, and Field C), you need to define the return type for each function.

After calculations are complete, you can exclude specific fields from the output using the Remove option. This feature helps eliminate temporary values that aren't needed in your final pipeline.

The Calculator delivers significantly faster execution compared to custom JavaScript scripts.
{% endhint %}

x

x

x

x
{% endtab %}

{% tab title="Group By" %}
{% hint style="info" %}
**Group by**

The Group By transform organizes rows from a data source according to one or more specified fields, creating a single row for each distinct group. Additionally, it can compute aggregate values like sums, averages, or counts for each group.

Typical applications include determining average sales figures by product category or tallying inventory quantities for each item type.

This step requires sorted input data to function properly. When working with unsorted data, only identical consecutive rows will be grouped together correctly. Furthermore, if data is sorted externally before entering the Transformation, differences in case sensitivity within the grouping fields may lead to unexpected results.

For scenarios involving unsorted input data, consider using the Memory Group By transform instead, which can handle data regardless of its initial order.
{% endhint %}

x

x

x
{% endtab %}

{% tab title="Text file output" %}
{% hint style="info" %}
**Text file output**

The Text File Output transform exports data to text file formats, most commonly generating CSV files that can be opened in spreadsheet applications like Excel.

This transform also supports creating fixed-width files by specifying field lengths in the fields configuration tab. You have two options for defining the output structure: use an existing Schema Definition or manually configure the field layout. When working with a Schema Definition, pair this transform with the Schema Mapping transform to align your incoming data stream with the chosen schema structure.
{% endhint %}

x

x

x
{% endtab %}
{% endtabs %}

x
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}


# Enrich Data

Enhance the quality of the data ..

{% hint style="info" %}

#### **Overview**

Data Enrichment is a value adding process, where external data from multiple sources is added to the existing data set to enhance the quality and richness of the data. This process provides more information of the product / service to the customer.

A common data enrichment process could, for example, correct likely misspellings or typographical errors in a database using precision algorithms. Following this logic, data enrichment could also add information to simple data tables.

Another way that data enrichment can work is in extrapolating data. Through methodologies such as fuzzy logic, engineers can produce more from a given raw data set. This and other projects can be described as data enrichment activities.

There are numerous data enhancement options available including:

* Telephone & Fax numbers
* Additional contact names
* Residential or Business location
* Standard Industrial Classification Codes (SIC’s) & ‘Market Sector’ codes
* No. of employees
* Small office /Home office (SoHo’s)
* Household income /Age
* Credit score
* Financial information
* Adding valuable geographic information and mapping (GIS) such as location analysis, distance calculations, spatial analysis, natural boundary analysis, and more
* Enhancing data by classifying, segmenting, and aggregating customer data using advanced statistical methodologies such as factor, cluster, and conjoint analysis.
  {% endhint %}

***


# Merge

When you merge rows and streams check the number of fields, data types and order.

{% hint style="info" %}
**Introduction**

In Pentaho Data Integration (PDI), true record merging differs from joining and focuses on combining or consolidating duplicate records into single entries:

The Append operation simply stacks records from two input streams. All rows from both streams appear in the output without any sorting or matching logic applied.

With Append, the output contains all records from the first stream followed immediately by all records from the second stream. Both input streams must share the same structure with compatible field types.
{% endhint %}

<figure><img src="/files/61vAQb6U0mvaKKUHBdk5" alt=""><figcaption><p>Merge streams</p></figcaption></figure>

{% hint style="info" %}
The Sorted Merge operation interleaves records from both input streams based on a predetermined sort order. This creates an integrated output where records are organized by their values.

For Sorted Merge to work properly, both input streams must be pre-sorted on the same field(s) before reaching the merge step. The operation preserves all records while maintaining the specified sort order.

Unlike joining operations, neither of these merging methods matches records based on key fields. They simply combine complete datasets according to different organizing principles - stacking for Append and interleaving by sort order for Sorted Merge.

Both techniques are valuable when you need to process records from multiple sources while maintaining all original data points.
{% endhint %}

<figure><img src="/files/BhvLWu30NspEVJ3d6mun" alt=""><figcaption><p>Sorted Merge</p></figcaption></figure>

***

{% hint style="info" %}
**Workshops**

The Dummy step in Pentaho Data Integration is a simple "do nothing" transformation that passes data through unchanged. It serves as a placeholder, helps join multiple streams, creates empty data rows when needed, and improves transformation organization.

The Merge Rows step compares two input data streams with identical structures to identify differences between them. It requires configuration of reference and compare streams, key fields for matching rows, and value fields to compare. The step outputs a single stream with all rows plus a "flagfield" indicating if each row is identical, changed, new, or deleted. This functionality is particularly useful for change data capture, data synchronization, audit trails, and implementing slowly changing dimensions.
{% endhint %}

{% tabs %}
{% tab title="Merge stream" %}
{% hint style="info" %}
**Merge stream - Dummy**

The Transformation underlines the ‘rules’ for manipulating data streams. Each data stream must have the same structure / layout, before they can be merged.

In this guided demonstration, you will merge data streams based on a set of rules:

• Add constant step
{% endhint %}

<figure><img src="/files/MZ7h0n4NplYDleSK9XY8" alt=""><figcaption><p>Merge streams</p></figcaption></figure>

{% content-ref url="/pages/TITwLgS97OVKMAbTSxKD" %}
[Merge Streams](/pentaho-data-integration/data-integration/enrich-data/merge/merge-streams)
{% endcontent-ref %}
{% endtab %}

{% tab title="Merge Rows (diff)" %}
{% hint style="info" %}
**Merge rows (diff)**

The Merge Rows (diff) compares the values between the merging rows and sets a ‘flag’.

In this guided demonstration, you will compare incoming records with reference records and then determine whether the record is Identical or needs updating, inserting, deleting:

• Merge Rows (diff) stream

• Merge Rows (diff) database
{% endhint %}

<figure><img src="/files/8mJPfTzw9HQLWoKu13XO" alt=""><figcaption><p>Merge Rows (diff)</p></figcaption></figure>

{% content-ref url="/pages/1jk5exCMFFTeAUXWCl7s" %}
[Merge Rows (diff)](/pentaho-data-integration/data-integration/enrich-data/merge/merge-rows-diff)
{% endcontent-ref %}
{% endtab %}
{% endtabs %}


# Merge Streams

{% hint style="warning" %}
**Workshop - Merge data streams**

The transformation underlines the ‘rules’ for manipulating data streams. Each data stream must have the same data stream fields / order / data type, before they can be merged.

In this workshop, you will need to add a 'description' to the data stream:

* Add constant step
  {% endhint %}

<figure><img src="/files/ctuPWS8RQ5QMOkArxB1s" alt="" width="563"><figcaption><p>Merge data streams</p></figcaption></figure>

***

{% tabs %}
{% tab title="English" %}

<figure><img src="/files/6aEwXgCNdgF0EUjZxefz" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="1. Text File Input" %}
{% hint style="info" %}
**Text File input**

The Text File Input step is used to read data from a variety of different text-file types. The most commonly used formats include Comma Separated Values (CSV files) generated by spreadsheets and fixed width flat files.

The Text File Input step provides you with the ability to specify a list of files to read, or a list of directories with wild cards in the form of regular expressions. In addition, you can accept filenames from a previous step making filename handling more even more generic.
{% endhint %}

1. Examine both the orders.txt and description.txt.
2. Configure the Text file input steps to point to, and retrieve the data from each of the files.
   {% endtab %}

{% tab title="2. Add Constant" %}
{% hint style="info" %}
**Add Constant**

The Add constant values step is a simple and high performance way to add constant values to the stream.

Why - in order to merge the streams each stream has to have the same layout.
{% endhint %}

1. The PRODUCTDESCRIPTION field is added to the ‘orders stream’ to ensure the data stream matches the ‘description’ data stream.
   {% endtab %}

{% tab title="3. Select Values" %}
{% hint style="info" %}
**Select Values**

The Select Values step is useful for selecting, removing, renaming, changing data types and configuring the length and precision of the fields on the stream. These operations are organized into different categories:

* Select and Alter - Specify the exact order and name in which the fields could be placed in the output rows
* Remove - Specify the fields that could be removed from the output rows
* Meta-data - Change the name, type, length and precision (the metadata) of one or more fields
  {% endhint %}

{% hint style="info" %}
Each of the Select values ensures that each data stream is consistent in layout before merging. Each field must be in the correct order within the data stream so that mappings are successful.
{% endhint %}

<figure><img src="/files/R8W7vPmDwKabmYM3a15T" alt=""><figcaption><p>Select values</p></figcaption></figure>
{% endtab %}

{% tab title="4. RUN" %}
{% hint style="info" %}
**RUN**
{% endhint %}

1. Run the Transformation.
2. Click on the Dummy step and ‘Preview’.

<figure><img src="/files/26UTcb20W5cBXhLWaHx6" alt=""><figcaption><p>Preview data</p></figcaption></figure>

{% hint style="info" %}
As you can see we have 2 merged streams ..
{% endhint %}
{% endtab %}
{% endtabs %}


# Merge Rows (diff)

Compare merging records ..

{% hint style="warning" %}
**Workshop - Merge Rows (diff)**

The Merge row (diff) compares the values between the merging rows and sets a ‘flag’.

In this workshop, you compare incoming records with reference 'golden' records to determine whether the record is Identical requires updating, inserting, or deleting:

* Merge rows (diff) stream
* Merge rows (diff) database
  {% endhint %}

{% embed url="<https://www.loom.com/share/cd56088673de462998803583425d0488?hideEmbedTopBar=true&hide_owner=true&hide_share=true&hide_title=true>" %}
Data Change Capture Sync after Merge
{% endembed %}

***

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

<figure><img src="/files/fb75f3FPphk5KdgXHdE5" alt="" width="563"><figcaption><p>Merge row (diff)</p></figcaption></figure>

{% tabs %}
{% tab title="1. Merged row (diff)" %}
{% hint style="info" %}
**Merge Rows (diff)**

Let's say we're doing a delta load of new data at specific times ..

Based on keys for comparison, we can use this step to merge reference rows (previous data) with compare rows (new data) to create merged output rows.

A flag in the row indicates how the values were compared and merged. Flag values include:

* identical

The key was found in both rows, and the compared values are identical.

* changed

The key was found in both rows, but one or more compared values are different.

* new

The key was not found in the reference rows.

* deleted

The key was not found in the compare rows.

If the rows are flagged as `deleted`, the merged output rows are created based upon the original reference rows stream.

For `identical`, `new`, or `changed` rows, the merged output rows are created based upon the original compare rows stream.
{% endhint %}

<figure><img src="/files/Qduc4moGoS9qNaZabVw0" alt="" width="563"><figcaption><p>Merge rows (diff)</p></figcaption></figure>
{% endtab %}

{% tab title="2. Synchronize after merge" %}
{% hint style="info" %}
**Synchronize after merge**

This step can be used in conjunction with the Merge Rows (diff) transformation step. The Merge Rows (diff) transformation step appends a Flag column to each row, with a value of "identical", "changed", "new" or "deleted".

This flag column is then used by the Synchronize after merge transformation step to carry out updates/inserts/deletes on a connection table.

This step uses the flag value to perform the sync operations on the database table.
{% endhint %}

<figure><img src="/files/SPnAe3weFzKzlZjO84s1" alt="" width="563"><figcaption><p>STG_ORDERS_MERGED</p></figcaption></figure>

{% hint style="info" %}

* Set the Key from both the Table and Stream.
* Get the Table / Stream Fields and ensure mapping is correct.
* <mark style="color:red;">Dont</mark> Update the Keys..!!
  {% endhint %}

<figure><img src="/files/BHZmgXWll1IWO0BisH72" alt="" width="563"><figcaption><p>Synchronize after merge - Advanced tab</p></figcaption></figure>

<table><thead><tr><th width="154.66666666666666">Option</th><th width="421">Description</th><th>Default Value</th></tr></thead><tbody><tr><td>Operation fieldname</td><td>This is a required field. This field is used by the step to obtain an operation flag for the current row.</td><td>flagfield</td></tr><tr><td>Insert when value equal</td><td>Specify the value of the Operation fieldname which signifies that anInsert should be carried out.</td><td>new</td></tr><tr><td>Update when value equal</td><td>Specify the value of the Operation fieldname which signifies that an Update should be carried out.</td><td>changed</td></tr><tr><td>Delete when value equal</td><td>Specify the value of the Operation fieldname which signifies that a Delete should be carried out.</td><td>deleted</td></tr><tr><td>Perform lookup</td><td>Performs a lookup when deleting or updating. If the lookup field is not found, then an exception is thrown. This option can be used as an extra check if you wish to check updates/deletes prior to their execution.</td><td></td></tr></tbody></table>
{% endtab %}

{% tab title="3. RUN" %}
{% hint style="info" %}
**RUN**

This step is aimed at reporting data marts .. delta loads to update the cube. Check out which records have undergone CRUID operations.
{% endhint %}

1. View the data in the Table.

<figure><img src="/files/ujbYWLiZXdD0B76QS6jn" alt="" width="563"><figcaption><p>STG_ORDERS_MERGED</p></figcaption></figure>

2. Run the Transformation with the hop between the Merge Rows (diff) and Synchronize after merge .. disabled.

<figure><img src="/files/pAcWoZx3WDy26yuhlenn" alt=""><figcaption><p>Synchronize after merge - FLAG</p></figcaption></figure>

3. Run the Transformation with the hop enabled.
4. Examine and compare the records.

<figure><img src="/files/aSsbWAUoB5ca3D6hxCea" alt="" width="563"><figcaption><p>STG_ORDERS_MERGED - Synchronize</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Joins

Pentaho Joins ..

{% hint style="info" %}
**Introduction**

Pentaho Data Integration (PDI) offers several join components to combine data from different streams based on specified key fields. Here's a summary of the main join types available:

**Merge Join**: This is the standard join step that combines two sorted input streams based on matching key fields. It supports inner joins, left outer joins, right outer joins, and full outer joins. Both input streams must be sorted on the join keys for this step to work correctly.

**Cross Join (Cartesian Product)**: This join creates all possible combinations of rows from two streams (a Cartesian product). It can be filtered to function as other join types by adding conditions. It's memory-intensive but doesn't require pre-sorted input.

**Database Join**: This specialized join allows you to look up values in a database table for each input row. It performs a database query for each incoming row, using values from the input stream as parameters.

**Multiway Merge Join**: This advanced join can combine more than two streams in a single operation, allowing for complex data integration scenarios when you need to merge multiple datasets together.

**XML Join**: A specialized join for combining XML data structures. It allows you to merge XML content from two streams, useful when working with XML-based data sources or targets.

Each join type has specific use cases and performance characteristics, allowing PDI to handle a wide variety of data integration scenarios efficiently.
{% endhint %}

<figure><img src="/files/LfKG4yZrKKvA6GMagJ0l" alt=""><figcaption><p>SQL joins</p></figcaption></figure>

***

{% hint style="info" %}
**Workshops**

There are different types of joins that you can use in Pentaho to combine data from different sources based on a common key or condition.

Here are some joins we are going to cover in this section:
{% endhint %}

{% tabs %}
{% tab title="Cross" %}
{% hint style="info" %}
**Cross Join**

A Pentaho cross join is a way of combining two streams of data in a Cartesian product, meaning that every row from one stream is joined with every row from the other stream. This can be useful for creating combinations of values or performing calculations based on multiple inputs.

However, a cross join can also result in a very large output, especially if the input streams have many rows. Therefore, it is important to optimize the cross join step by using filters, conditions, or lookups to reduce the number of rows or columns in the output.
{% endhint %}

<figure><img src="/files/Es830VUvgevh8Vi7rWhF" alt=""><figcaption><p>Cross Join</p></figcaption></figure>

{% content-ref url="/pages/DfgQQioHIyeSAOFCtAAJ" %}
[Cross Join](/pentaho-data-integration/data-integration/enrich-data/joins/cross-join)
{% endcontent-ref %}
{% endtab %}

{% tab title="Merge" %}
{% hint style="info" %}
**Merge Join**

There are four basic types of SQL joins: inner, left, right, and full. The easiest and most intuitive way to explain the difference between these four types is by using a Venn diagram, which shows all possible logical relations between data sets.\
Again, it's important to stress that before you can begin using any join type, you'll need to extract the data and load it into an RDBMS, where you can query tables from multiple sources.
{% endhint %}

<figure><img src="/files/zgymkeklj3l1Md6hrR2F" alt=""><figcaption><p>Joins</p></figcaption></figure>

{% content-ref url="/pages/dDM5Ae2aQtvRpRjU61lw" %}
[Merge Join](/pentaho-data-integration/data-integration/enrich-data/joins/merge-join)
{% endcontent-ref %}
{% endtab %}

{% tab title="Database" %}
{% hint style="info" %}
**Database**

Database Join is a powerful step in Pentaho Data Integration that allows you to enhance your data stream with information from a database using SQL queries.

Unlike regular join steps that operate on two data streams, the Database Join connects your transformation's data stream directly to a database. It uses values from the incoming stream as parameters in SQL queries.

For each row in your input stream, PDI executes a parameterized SQL query against the specified database connection. The query results are then added as new fields to the original row.

This step is particularly efficient when you need to look up relatively small amounts of data from a database based on values in your stream. It leverages database optimization rather than performing joins within PDI's memory.

The Database Join requires a valid database connection and a properly formatted SQL query that references input fields as parameters (typically using ? placeholders).

One key limitation is that each row triggers a separate database query, which can cause performance issues with large input streams. For better performance with significant data volumes, consider using the Table Input step with a cached connection.

Be careful with the SQL query complexity, as overly complex queries may impact transformation performance. The step works best for simple lookups rather than complex analytical queries.
{% endhint %}

<figure><img src="/files/KNuBBnsXuTLVOt6rywHx" alt=""><figcaption><p>Database Join</p></figcaption></figure>

{% content-ref url="/pages/VMFOmq5edDd75vbQtFnc" %}
[Database Join](/pentaho-data-integration/data-integration/enrich-data/joins/database-join)
{% endcontent-ref %}
{% endtab %}

{% tab title="XML" %}

{% endtab %}
{% endtabs %}


# Cross Join

Good old Cartesian Join ..

{% hint style="warning" %}
**Workshop - Cross Join**

The CARTESIAN JOIN or CROSS JOIN returns the Cartesian product of the sets of records from two or more joined tables. Thus, it equates to an inner join where the join-condition always evaluates to either True or where the join-condition is absent from the statement .. whatever that means .. basically, its every possible combination.

In this workshop we'll be cross joining first names with middle names and again with our surname.
{% endhint %}

<figure><img src="/files/1RLemhlsRLR0IKoWTJL2" alt=""><figcaption><p>Cross / Cartesian Join</p></figcaption></figure>

<figure><img src="/files/lJ6SBvIQXEXu3MsOGf7P" alt=""><figcaption><p>Cross Joins</p></figcaption></figure>

***

{% tabs %}
{% tab title="English" %}

<figure><img src="/files/H7ajXCDy7BysZd2TeNvb" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="1. Data Grid" %}
{% hint style="info" %}
You should be familiar with the Data Grid step. Used to list values for first and middle names.
{% endhint %}

<div><figure><img src="/files/YiVVG6w6THmDgleYMZLA" alt=""><figcaption><p>First names</p></figcaption></figure> <figure><img src="/files/Wz3KnlV3tO0WDYh09y24" alt=""><figcaption><p>Middle names</p></figcaption></figure></div>
{% endtab %}

{% tab title="2. Join" %}
{% hint style="info" %}
Joins all possible first\_name, middle\_name combinations together.
{% endhint %}

<figure><img src="/files/YQaHYTKpxLDxfX6o3nQc" alt="" width="375"><figcaption><p>Join rows</p></figcaption></figure>

{% hint style="info" %}
You can also add a condition to constrain the resulting dataset.
{% endhint %}
{% endtab %}

{% tab title="3. UDJE" %}
{% hint style="info" %}
2 new fields are added to the data stream:

* first\_middle\_name: concates the first\_names and middle names.
* initials: returns the 2 character initials.

This method has two variants and returns a new string that is a substring of this string. The substring begins with the character at the specified index and extends to the end of this string or up to endIndex – 1, if the second argument is given.
{% endhint %}

<figure><img src="/files/K7IQXzxKL3FKjF58cSEu" alt="" width="563"><figcaption><p>UDJE - concat names, initials</p></figcaption></figure>
{% endtab %}

{% tab title="4. Get variable" %}
{% hint style="info" %}
Returns the value associated with the ${surname}. This is set in the Parameters tab in Transformation properties.
{% endhint %}

<figure><img src="/files/fqUMIiISnKNL0qMdS6zF" alt="" width="563"><figcaption></figcaption></figure>
{% endtab %}

{% tab title="5. Join" %}
{% hint style="info" %}
Joins all possible first\_middle\_name, surname combinations together. The output for initials is also excludes in the list various initials combinations.
{% endhint %}

<figure><img src="/files/wau5duwmXi7J7a4dmjMq" alt="" width="563"><figcaption><p>Join rows</p></figcaption></figure>
{% endtab %}

{% tab title="6. Select values" %}
{% hint style="info" %}
Determine the order and selection of the data stream fields.
{% endhint %}

<figure><img src="/files/NQRozBNCuyEXDK2LWhPV" alt="" width="563"><figcaption><p>Select values</p></figcaption></figure>
{% endtab %}

{% tab title="7. UDJE" %}
{% hint style="info" %}
2 new fields are added to the data stream:

* boys\_initials: returns the babys’ 3 character initials.
* boys\_name: concates first\_name + middle\_name + surname
  {% endhint %}

<figure><img src="/files/w2imlH2Bm8OQDna7CeKd" alt=""><figcaption><p>UDJE - name</p></figcaption></figure>
{% endtab %}

{% tab title="8. Reservoir Sampling" %}
{% hint style="info" %}

* Returns 5 sampled records.
* Returns Random seed ${seed}
* Reservoir Sampling allows you to select a set number of random records, from an unknown number ‘reservoir’ of records, i.e. not known beforehand.
* Use a different seed value to ensure no two ‘sets’ are the same.
  {% endhint %}

<figure><img src="/files/JGVeho97eBsHIpMytPzp" alt="" width="375"><figcaption><p>Reservoir sampling</p></figcaption></figure>
{% endtab %}

{% tab title="9. RUN" %}
{% hint style="info" %}
**RUN**

The workshop illustrates the use of cross joins to create data sets with every possible combination - unless conditions are set. The final dataset is randomly selected using Reservoir Sampling - a common technique used in ML.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Click on the Preview tab:

<figure><img src="/files/A1P9cik5lGRWoTnUXxJ8" alt=""><figcaption><p>Preview data</p></figcaption></figure>
{% endtab %}
{% endtabs %}


# Merge Join

Standard SQL joins ..

{% hint style="warning" %}
**Workshop - Merge Join**

A workshop to illustrate various SQL joins.

In this workshop, we're going to run through the various join types available in the Merge join step.
{% endhint %}

{% embed url="<https://www.loom.com/share/cdbbff783c7549e09cf9fddef5ccefde>" %}

***

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

<figure><img src="/files/dgHWlaPxq9BeHE4bFoXn" alt="" width="375"><figcaption><p>Merge Join</p></figcaption></figure>

{% tabs %}
{% tab title="1. Data Grid" %}
{% hint style="info" %}
**Data grid**

The Data grid step allows you to enter a static list of rows in a grid. This is usually done for testing, reference or demo purposes.

Options

* Meta tab: You can specify the field metadata (output specification) of the data
* Data tab: This grid contains the data. Everything is entered in String format so make sure you use the correct format masks in the metadata tab.
  {% endhint %}

1. Drag the Data Grid step onto the canvas.
2. Open the Data Grid properties dialog box.
3. Ensure the following details are configured, as outlined below:

<div align="left"><figure><img src="/files/nXZbbLNpCsFBPVbaz5ib" alt=""><figcaption><p>ABCD - Data Grid</p></figcaption></figure> <figure><img src="/files/5h0mHmmKTNOQMoM8rYxo" alt=""><figcaption><p>Red Blue Yellow - data grid</p></figcaption></figure></div>
{% endtab %}

{% tab title="2. Merge Join" %}
{% hint style="info" %}
**Merge Join**

The Merge Join step performs a classic merge join between data sets with data coming from two different input steps. Join options include INNER, LEFT OUTER, RIGHT OUTER, and FULL OUTER.
{% endhint %}

1. Drag the Merge Join step onto the canvas.
2. Open the Merge Join properties dialog box.
3. Select various Join Types to view the resulting dataset

<figure><img src="/files/xvgn45La24cNjvq0mqex" alt="" width="375"><figcaption><p>Merge Join</p></figcaption></figure>

{% hint style="warning" %}
Obviously you need to join on a unique key(s)
{% endhint %}
{% endtab %}

{% tab title="3. RUN" %}

1. Click the Run button in the Canvas Toolbar.
2. Click on the Dummy step Preview tab:

**INNER Join**

<figure><img src="/files/JLWjg6Mzj00L9mAGcQru" alt=""><figcaption><p>INNER Join</p></figcaption></figure>

**LEFT OUTER Join**

<figure><img src="/files/2xYxox7DsedbQ0xUiTfa" alt=""><figcaption><p>LEFT OUTER Join</p></figcaption></figure>

**RIGHT OUTER Join**

<figure><img src="/files/sgFXkfYOdmGMpW2YUZuB" alt=""><figcaption><p>RIGHT OUTER Join</p></figcaption></figure>

**FULL OUTER Join**

<figure><img src="/files/kyFju85hBemtlZ0Fa0nu" alt=""><figcaption><p>FULL OUTER Join</p></figcaption></figure>

{% hint style="info" %}
Now give it a go with the 'Merge Streams' scenario ..
{% endhint %}
{% endtab %}
{% endtabs %}


# Database Join

A self join or recursive join ..  or is it ?

{% hint style="warning" %}
**Workshop - Database Join**

Searching for information in databases, text files, web services, and so on, is a very common task. In this workshop we're going to query the Products table for products are listed below a set buy price.

The database join isn't actually a join, but a series of queries against the table based on set conditions. Be aware this results in a performance hit.
{% endhint %}

<figure><img src="/files/EYLxFSiDXVkEU6zAdDUO" alt=""><figcaption><p>Database Join</p></figcaption></figure>

***

{% tabs %}
{% tab title="English" %}

<figure><img src="/files/RovcatI7zyAjpU44AmZu" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="1. Data Grid" %}
{% hint style="info" %}
**Data grid**

The Data grid step allows you to enter a static list of rows in a grid. This is usually done for testing, reference or demo purposes.
{% endhint %}

1. Drag the Data grid step onto the canvas.
2. Open the Data grid properties dialog box.
3. Ensure the following details are configured, as outlined below:

<div><figure><img src="/files/Nfb4zIz3eUc77ukvCx4u" alt=""><figcaption><p>Data grid - Meta</p></figcaption></figure> <figure><img src="/files/YuRRhTSKMEDiNZ4StKkw" alt=""><figcaption><p>Data grid - Data</p></figcaption></figure></div>
{% endtab %}

{% tab title="2. Database Join" %}
{% hint style="info" %}
**Database Join**

The Database Join step allows you to run a query against a database using data obtained from previous steps. The parameters for this query are specified as follows:

* The data grid in the step properties dialog. This allows you to select the data coming in from the source hop.
* As question marks (?) in the SQL query. When the step runs, these will be replaced with data coming in from the fields defined from the data grid. The question marks will be replaced in the same order as defined in the data grid.
  {% endhint %}

1. Drag the Database Join step onto the canvas.
2. Open the Database Join properties dialog box.
3. Ensure the following details are configured, as outlined below:

<figure><img src="/files/c4gWc0yLytVC32xEKLf9" alt="" width="563"><figcaption><p>Database join</p></figcaption></figure>

{% hint style="info" %}
The ‘Parameter fieldname’ is where you specify the parameters, therefore the values, for the conditions. Each row in the grid represents a comparison between a column in the table, and a field in your stream, by using one of the provided comparators.

**LIKE** matches values. You can't alias a column in the select clause and then use it in the where clause

The question marks you type in the SQL statement represent parameters. The purpose of these parameters is to be replaced with the fields you provide in ‘Parameter fieldname’. For each row in the stream, the Database join step replaces the parameters in the same order as they are in the grid, and executes the SQL statement.

So, let’s look at the WHERE conditions entered:

PRODUCTNAME LIKE like\_statement and BUYPRICE < max\_price

For the first record this translates as:

WHERE PRODUCTNAME LIKE concat ('%','Aston Martin','%') AND BUYPRICE < 90

As the Outer Join option is checked The FULL OUTER JOIN keyword returns all rows from the left table and from the right table. The FULL OUTER JOIN keyword combines the result of both LEFT and RIGHT joins.

<img src="/files/k0nGugXMcitF9j6a7o9e" alt="" data-size="original">

The table dataset A is then compared with the stream dataset B. If there’s a match, then values for PRODUCTNAME and PRODUCTSCALE are returned.

*This is not a database join. Instead of joining tables in a database, you are joining the result of a database query with a dataset.*

For the second record:

WHERE PRODUCTNAME LIKE concat ('%','Ford Falcon','%') AND BUYPRICE < 70

As there is no record, NULL values are returned for:

PRODUCTNAME and PRODUCTSCALE.

So far, the results could be achieved using a Database Lookup step. However, there is a significant difference, as illustrated with the third row. For Corvette, the Database join found two matching rows in the database, and retrieved them both. Not possible with a Database lookup step.
{% endhint %}
{% endtab %}

{% tab title="3. RUN" %}
{% hint style="info" %}
**RUN**

A Database Join involves running a bunch of queries with condition against a table. Useful when you're expecting to return a few records.
{% endhint %}

1. Click the Run button in the Canvas Toolbar.
2. Click on the Preview tab:

<figure><img src="/files/wdB5VczqrC9oeVDQs7DL" alt=""><figcaption><p>Results</p></figcaption></figure>

{% hint style="info" %}
Note that there is more than one Corvette product. The database join is querying the table to return all the values, even NULL.
{% endhint %}
{% endtab %}
{% endtabs %}


# XML Join

Join XML streams ..

{% hint style="warning" %}
**Workshop - XML Join**
{% endhint %}

{% hint style="info" %}

#### XML Join in Pentaho Data Integration

XML Join is a specialized step in Pentaho Data Integration designed to incorporate XML content into your data stream based on values from another stream.

This step accepts two input streams - the main data stream and an XML stream. It merges them by adding the XML content as a new field in your main data stream.

The XML stream must contain well-formed XML data that will be integrated into your transformation. The main stream contains the records you want to enhance with this XML content.

For each row in the main stream, PDI matches it with corresponding XML content based on a specified join key. The XML content is then added as a new field to the main stream row.

XML Join is particularly useful when dealing with web services, XML databases, or when you need to construct complex XML documents from relational data sources.

The step offers options to specify the target XML field name, the join comparison field, and the ability to encode the XML content if needed for further processing.

When configuring XML Join, you must specify which stream provides the XML content and which one serves as the main stream. The order of connecting these streams to the step is critical to its proper functioning.
{% endhint %}

***

{% tabs %}
{% tab title="English" %}

<figure><img src="/files/z7vFACJBMtrgJuTTeKXt" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="First Tab" %}

{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}


# Lookups

{% hint style="info" %}
**Introduction**

Besides transforming the data, you may need to search and bring data from other sources. Let us look at the following examples:

* You have some product codes and you want to look for their descriptions in an Excel file
* You have a value and want to get all products whose price is below that value from a database

Searching for information in databases, text files, web services, and so on, is a very common task, and Kettle has several steps for doing it.
{% endhint %}

<figure><img src="/files/FHHpcOeIcuhNrFJJx0rk" alt=""><figcaption><p>Lookup / List tables</p></figcaption></figure>

{% hint style="info" %}
For instance, if your database is about sales, you probably have a Customers table and an Orders table, each with its own attributes resolved through a Foreign Key. The lookup tables are usually very small, with just a handful of rows in them.
{% endhint %}

***


# Database Lookups

{% hint style="warning" %}
**Workshop - Database Lookup**

The Database lookup step allows you to look for values in a database table.

In this guided demonstration, you will:

* Configure the following steps:
  * Database Lookup
    {% endhint %}

<figure><img src="/files/B6caO2FHf9CdGYoIys1d" alt="" width="563"><figcaption><p>Database lookup</p></figcaption></figure>

***

{% tabs %}
{% tab title="First Tab" %}

<figure><img src="/files/22tlX0Fee58hC8wtZlBC" alt=""><figcaption></figcaption></figure>
{% endtab %}

{% tab title="Second Tab" %}

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Create a new Transformation**

Any one of these actions opens a new Transformation tab for you to begin designing your transformation.

* By clicking File > New > Transformation
* By using the CTRL-N hot key
  {% endhint %}

{% tabs %}
{% tab title="1. Data Grid" %}
{% hint style="info" %}
**Data grid**

The Data Grid step allows you to enter a static list of rows in a grid. This is usually done for testing, reference or demo purposes.
{% endhint %}

1. Drag the Data Grid step onto the canvas.
2. Open the Data grid properties dialog box.
3. Ensure the following details are configured, as outlined below:

<div><figure><img src="/files/Nfb4zIz3eUc77ukvCx4u" alt=""><figcaption><p>Data grid - Meta</p></figcaption></figure> <figure><img src="/files/YuRRhTSKMEDiNZ4StKkw" alt=""><figcaption><p>Data grid - Data</p></figcaption></figure></div>
{% endtab %}

{% tab title="2. UDJE" %}
{% hint style="info" %}
**User defined Java expression**

The User Defined Java Expression step in Pentaho Data Integration allows you to write custom Java code that executes on each row of your data transformation. This step is useful when you need to perform complex calculations or data manipulations that aren't possible with PDI's standard steps.

You can access field values using the `get("fieldname")` method and create multiple expressions within a single step. Each expression produces a new output field in your data stream. The step handles type conversion automatically, making it flexible for various data operations.

Common uses include mathematical calculations, string manipulations, conditional logic, and date transformations. It's particularly valuable when you need Java-specific functionality or want to simplify your transformation by replacing multiple basic steps with a single, powerful Java expression.
{% endhint %}

1. Drag the User Defined Java expression step onto the canvas.
2. Open the User defined Java expression properties dialog box.
3. Ensure the following details are configured, as outlined below:

<figure><img src="/files/l1fKrBcvuIsUMCTWDIV4" alt="" width="563"><figcaption><p>UDJE - like statement</p></figcaption></figure>

{% hint style="info" %}
The LIKE operator is used in a WHERE clause to search for a specified pattern in a column.
{% endhint %}
{% endtab %}

{% tab title="3. Database Lookup" %}
{% hint style="info" %}
**Database lookup**

The Database lookup step has 3 options
{% endhint %}

{% tabs %}
{% tab title="3.1 Simple" %}
{% hint style="info" %}
**Simple Lookup**

The Database lookup step allows you to look up values in a database table. Lookup values are added as new fields onto the stream.
{% endhint %}

1. Drag the Data Grid step onto the canvas.
2. Open the Data grid properties dialog box.
3. Ensure the following details are configured, as outlined below:

<figure><img src="/files/tC27igd9K9GVJJjrVEYY" alt=""><figcaption><p>Dtabase lookup - simple</p></figcaption></figure>

{% hint style="info" %}
The ‘key’ fields are where you specify the conditions. Each row in the grid represents a comparison between a column in the table, and a field in your stream, by using one of the provided comparators.

In this example:

WHERE PRODUCTNAME LIKE '%Aston Martin%' AND BUYPRICE < 90 WHERE PRODUCTNAME LIKE '%'Ford Falcon%' AND BUYPRICE < 70 WHERE PRODUCTNAME LIKE '%Corvette'%' AND BUYPRICE < 70
{% endhint %}

{% hint style="info" %}
The Database lookup step allow us to retrieve any number of columns based on the search criteria. Each database column you enter in the lower grid will become a new field in your dataset.

You can rename them (this is particularly useful if you already have a field with the same name) and supply a default value if no record is found in the search. In the workflow, you added three fields: PRODUCTNAME, PRODUCTSCALE, and BUYPRICE.

For values where there’s no match for PRODUCTNAME, ‘not available’ is returned. In the Preview, notice there are no PRODUCTNAMES that match %Ford Falcon% where the max price is less than the max price of 70.
{% endhint %}
{% endtab %}

{% tab title="3.2 Error Handling" %}
{% hint style="info" %}
In this workflow, error handling has been enabled, with a write to log step.
{% endhint %}

1. To see this in action, disable the Hops to Database Lookup (simple) and Database Lookup (do not pass).
2. The error message is written out in the Logging output.

<figure><img src="/files/qcEAQQ7vZM1er9pQGPWk" alt=""><figcaption><p>Logging Results</p></figcaption></figure>

3. Preview the Database Lookup (with error handling) step.

<figure><img src="/files/hTMi9Ac7s9OgH0kI3QYc" alt=""><figcaption><p>Database lookup - erorr handling</p></figcaption></figure>

{% hint style="info" %}
The rows for which the lookup fails, go directly to the stream that captures the error, in this case, the ‘Write to log’ step.
{% endhint %}
{% endtab %}

{% tab title="3.3 Do not pass" %}
{% hint style="info" %}
Taking some action when there are too many results The Database lookup step is meant to retrieve just one row of the table for each row in your dataset. If the search finds more than one row, the following two things may happen:

1. If you check the Fail on multiple results? option, the rows for which the lookup retrieves more than one row will cause the step to fail. In that case, in the Logging tab window, you will see an error similar to the following: ...

\- Database lookup (fail on multiple res.).0 – ERROR... Because of an error, this step can't continue:

\- Database lookup (fail on multiple res.).0 – ERROR: Only 1 row was expected as a result of a lookup, and at least 2 were found! Then you can decide whether you want to leave the transformation or capture the error.

2. If you don't check the Fail on multiple results? option, the step will return the first row it encounters. You can decide which one to return by specifying the order. You do that by typing an order clause in the Order by textbox. In the Sampledata database, there are three products that meet the conditions for the Corvette row. If, for Order by, you type PRODUCTSCALE DESC, PRODUCTNAME, then you will get 1958 Chevy Corvette Limited Edition, which is the first product after ordering the three found products by the specified criterion.

If, instead of taking some of those actions, you realize that you need all the resulting rows, you should take another approach—replace the Database lookup step with a Database join or a Dynamic SQL row step.

Compare this with the Database Join

As the database join is a full outer, all the records are returned from the database table, rather than just return a single lookup reference value.
{% endhint %}
{% endtab %}
{% endtabs %}
{% endtab %}
{% endtabs %}




---

[Next Page](/llms-full.txt/1)

