Intro
Apache Hop (Hop Orchestration Platform) is an open-source data orchestration and engineering platform built on the Kettle codebase. It lets teams design, schedule, and monitor ETL pipelines and workflows visually. This article addresses the advanced concepts that engineers need once they move beyond drag-and-drop basics: how Hop is structured, how pipelines and workflows behave, how metadata injection and variables work, how to configure logging and error handling, and how to operate Hop safely in production.
This guide combines a conceptual deep dive with concrete, copy-paste-ready examples. Each section explains an area of Hop's architecture and then shows how to verify, test, or change the relevant behavior. Every command includes explicit version checks, read-only observations before changes, and recovery steps. The focus is on safe operations: know your current state, limit the impact of changes, and verify outcomes.
Although Apache Airflow, NiFi, and similar tools are sometimes compared to Hop, this article mentions them only when they affect compatibility, security, or migration decisions.
Version and Environment Inventory
Why It Matters
Before making any change, you need a precise inventory of your Hop installation: the version, how it is deployed, and how pipelines are organized. This reduces the risk of applying documentation or commands meant for a different release. Hop's behavior can change between versions, especially around the REST API, plugin loading, and the hop-server component.
Read-Only Commands
Start with these read-only checks:
# 1. Hop version
hop-conf --version
# Expected output example: 2.1.0
# 2. Hop home directory
hop-conf --hop-home
# Expected output example: /opt/hop
# 3. Environment variables
hop-conf --environment
# Lists configured environments, such as development, production
If the hop-conf command is not on your PATH, run it from the Hop installation directory:
cd /opt/hop
./hop-conf --version
For installations running as a service, inspect the service definition:
# Linux (systemd)
systemctl cat hop
# Shows service file, including ExecStart and environment file
The output tells you which Java version Hop is using, memory settings, and the working directory. Note these values before any change.
Deployment Topology
Hop can run in three main modes:
- Local (fat client): You run
hop-guiorhop-rundirectly on a machine. Suitable for development. - Remote engine (hop-server): A long-running server executes pipelines and workflows submitted from a client or via REST.
- Clustered: Multiple Hop servers coordinated by a lightweight cluster manager (often ZooKeeper) for distributed execution.
Each mode has different prerequisites and failure modes. For example, a missing hop-server binary on the client means you cannot submit remote executions, but that binary is not needed if you only run locally.
Smallest Justified Change Example
Suppose you find that your production Hop server is running version 2.0.0, but the documented examples require 2.1.0 features. Before upgrading, capture the current state:
# Record installed plugins
hop-conf --plugins > hop_plugins_before.txt
# Record current server configuration
cp /opt/hop/config/hop-server.xml /opt/hop/config/hop-server.xml.backup
Then upgrade only the server binary, keeping the configuration backup. After upgrade, verify:
hop-conf --version
# Expected: 2.1.0
# Check that required plugins are still present
hop-conf --plugins | grep "Apache Hop Transform"
If the version check fails or a plugin is missing, restore from backup and review the upgrade procedure.
Safe Configuration Path
Core Configuration Files
Hop's configuration lives in several XML files:
config/hop-config.xml: Global properties, logging settings.config/hop-server.xml: Remote server settings (host, port, authentication).config/environments.xml: Defines environment-specific variables.config/projects.xml: Project home directories.
An advanced concept is environment separation. Never hardcode database credentials or file paths in pipelines. Instead, define them as variables in an environment and reference them with ${VARIABLE_NAME}.
Example: Environment-Safe Database Connection
Create a new environment called production:
hop-conf --environment-create production
Edit the environment file (typically config/environments/production.json) and add variables:
{
"DB_HOST": "db.internal.example",
"DB_PORT": "5432",
"DB_NAME": "warehouse",
"DB_USER": "etl_user",
"DB_PASSWORD": "replace_with_secret_placeholder"
}
In your pipeline, create a PostgreSQL connection with these settings:
- Host:
${DB_HOST} - Port:
${DB_PORT} - Database:
${DB_NAME} - Username:
${DB_USER} - Password:
${DB_PASSWORD}
When you run the pipeline, activate the environment:
hop-run -e production -f /path/to/pipeline.hpl
This prevents accidental deployment of development credentials to production. Always use a secrets manager or environment variable injection for actual passwords; do not commit the environment file to version control if it contains plaintext secrets.
Testing Configuration Changes
Before changing a production configuration file, test it on a copied instance. For example, to test a new logging level:
# Copy original config
cp config/hop-config.xml config/hop-config.xml.backup
# Change log level in copy (e.g., from INFO to DEBUG)
# Run Hop with the modified config
hop-run -c config/hop-config-test.xml -f /path/to/test_pipeline.hpl
# Check log output
hop-log-viewer
If the test behaves as expected, promote the configuration. Otherwise, restore the backup.
Verification and Diagnostics
Pipeline Metrics and Logs
Hop records execution metrics in its database or log files. To diagnose a slow transformation, you need to see metrics like rows read, written, and error counts per step.
Enable step metrics in your pipeline:
- Right-click on a transformation step and select Metrics.
- Choose the metrics you need (e.g.,
LinesInput,LinesOutput,LinesRejected,Errors). - Save and run the pipeline.
- After the run, open the Metrics tab in the Hop GUI to see per-step numbers.
Alternatively, use the command line to capture metrics:
hop-run -f /path/to/pipeline.hpl -l /path/to/metrics.log
The log file contains lines like:
2023/09/01 12:00:01 - Table output.0 - Finished processing (I=1000, O=1000, R=0, W=1000, U=0, E=0)
Here I, O, R, W, U, E stand for input, output, rejected, written, updated, and error rows. If E (errors) is greater than zero, inspect the error rows via the error handling step.
Using Hop Server for Remote Diagnostics
If pipelines run remotely, enable hop-server diagnostics:
hop-server -h localhost -p 8080 -u cluster -p secret -l /path/to/server.log
Then query the server via its REST API:
# Get server status
curl -u cluster:secret http://localhost:8080/hop/status/
# Expected output: JSON with status "running" and pipeline list
# Get last execution details
curl -u cluster:secret http://localhost:8080/hop/executions/?name=my_pipeline
Use these read-only endpoints before altering a running pipeline.
Diagnosing Common Errors
Error: Transformation cannot be found
Check the file path and extension. Hop expects .hpl for pipelines and .hwf for workflows.
ls -l /path/to/your_pipeline.hpl
# If file exists, check project references
hop-conf --projects
# Ensure the project points to the correct directory
Error: Database connection failed
Test connectivity using Hop's database explorer:
hop-db-explorer -j jdbc:postgresql://db.example:5432/warehouse -u etl_user -p secret -t "SELECT 1"
If this fails, check network access and credentials. Do not modify the pipeline's connection settings until you confirm the database is reachable.
Failure Modes and Recovery
Understanding Hop's Error Handling Architecture
Hop distinguishes between pipeline errors (a transform fails to process rows) and workflow errors (a job action fails). Understanding this is essential for designing recovery.
In a pipeline, you can attach an error handling step to any transform. When the transform encounters an error, the problematic rows are sent to the error handling step instead of halting the whole pipeline.
Example: A Table Input step reads from a database. We attach a Text File Output step as the error handler to write rejected rows.
Configuration:
- Select the Table Input step.
- In its properties, go to the Error handling tab.
- Set the target step to the Text File Output step.
- Optionally, set the error limit (e.g., stop after 100 errors).
- In the Text File Output step, define fields for error message, error code, and original row data.
When you run the pipeline, rows that fail database constraints or data type conversions are diverted to the text file. The pipeline continues processing other rows.
Workflow Error Handling
Workflows execute actions in sequence or parallel. By default, a failed action stops the workflow. To execute recovery actions, use Hop Action steps like Abort, Mail, or Write To Log.
Example workflow:
- Start -> Run Pipeline (ETL) -> Success: Send Mail -> Done
- If Run Pipeline fails -> Write To Log (error details) -> Abort
To implement conditional execution, right-click on the connection and set the Result condition (e.g., Success, Failure, Always).
Automatic Retry
Some actions support retry. For a database connection, you can set a retry count and wait interval on the action's properties. However, retries are not a substitute for proper error handling; always design for graceful failure and notification.
Recovery Commands
If a pipeline fails mid-run, recover by checking its log:
hop-log-viewer -f /path/to/pipeline.log
Identify the failing step and error message. Common fixes:
- Out of memory: Increase Java heap in
hop-config.xml(e.g.,-Xmx4g). - Locked database rows: Wait or kill the blocking session.
- Table already exists: Drop or rename the target table, or change the pipeline to use truncate or update.
Restart the pipeline after the fix. Use checkpointing for very large pipelines: Add a Checkpoint step between major sections. If the pipeline fails after the checkpoint, you can resume from the checkpoint files instead of restarting from zero.
Common Pitfalls and How to Avoid Them
1. Running with the Wrong Environment
Why it happens: Developers often have multiple environments (dev, test, prod). Forgetting to specify -e or selecting the wrong environment in the GUI leads to writing data to the wrong database.
How to avoid: In the Hop GUI, check the environment indicator in the top bar before running. In scripts, always pass -e <environment> explicitly. Consider using separate Hop installations or projects for each environment.
Recovery: Immediately stop the pipeline. Check the target database for rows inserted in the wrong environment. Use backup or transaction logs to remove or correct them if possible.
2. Ignoring Variable Scoping
Why it happens: Variables can be defined at system, environment, pipeline, and workflow levels. The precedence rules are not always obvious; a pipeline-level variable can override an environment variable unintentionally.
How to avoid: Document variables clearly and use a naming convention like ENV_DB_HOST. In the pipeline properties, review variables and parameters before running. Use the Set Variables step cautiously, as it modifies the variable space for the entire pipeline.
Recovery: If unexpected variables cause data corruption, stop the run and check the log for variable values at execution start. Correct the variable definition and rerun.
3. Overlooking Metadata Injection Complexity
Why it happens: Metadata injection is powerful but can produce confusing errors when the injected metadata does not match the target step's expectations. For example, injecting a field name that does not exist in the target step.
How to avoid: Start with a simple template step and inject only necessary fields. Test with a small dataset. Use the Metadata tab to view the actual metadata passed.
Recovery: If injection fails, examine the error log to see which metadata item caused the failure. Adjust the injection mapping and retry.
4. Neglecting Logging Configuration
Why it happens: Default logging may not capture enough detail for post-mortem analysis, or it may capture too much (sensitive data) and fill disk space.
How to avoid: Set appropriate log levels per package in hop-config.xml. For example, set org.apache.hop to INFO, but your own package to DEBUG. Use log rotation.
Recovery: If logs are insufficient, temporarily increase logging to DEBUG and rerun the failing pipeline. Be aware of performance impact and sensitive data exposure.
5. Modifying Production Pipelines Without Version Control
Why it happens: Hop pipelines are XML files. Without version control, changes are lost, and rollback is impossible.
How to avoid: Store all .hpl, .hwf, and environment files in a Git repository. Use branches for changes and require review before merging to production. Hop can integrate with Git via plugins or external scripts.
Recovery: If a bad change is deployed, checkout the previous version of the pipeline file and reload it into Hop.
Operations Checklist
Use this checklist before and after any operation on a Hop installation or pipeline. Assign a single accountable owner for each item; the owner should be an engineer who understands the specific area. Review this checklist at least monthly, or immediately after any incident.
| Item | Owner | Frequency | Verification Command | Expected Result | Recovery Action If Fail |
|---|---|---|---|---|---|
| Hop version is documented and matches across environments | Priya Shah, DevOps Lead | Monthly | hop-conf --version | Same version string (e.g., 2.1.0) on all nodes | Upgrade or downgrade to match, then test |
| Environment variables are correctly set for production | Alex Chen, Data Engineer | Weekly before production runs | hop-conf --environment --list and inspect active environment | Production environment active with correct variables | Stop run, correct environment, rerun |
| Database connections respond | Marcus Reid, DBA | Daily (automated) | hop-db-explorer -j <jdbc_url> -u <user> -p <secret> -t "SELECT 1" | Returns 1 | Investigate network, credentials, database status |
| Pipeline logs are rotated and not filling disk | Sofia Garcia, Ops Engineer | Weekly | df -h /path/to/logs | Disk usage below 80% | Clean old logs, adjust log rotation |
| Error handling steps attached to critical transforms | David Kim, Pipeline Developer | Per pipeline change | Inspect pipeline XML for error_handling elements | Each critical transform has error handling | Add error handling step, test |
| Hop server health check passes | Lena Fischer, Platform Engineer | Every 5 minutes | curl -u <user>:<secret> http://<server>:8080/hop/status/ | HTTP 200 with "status":"running" | Restart hop-server, check logs |
| Backups of Hop configuration and project files exist | Omar Hassan, Systems Admin | Daily | Check backup job logs or run ls -l /backup/hop/ | Today's backup files present | Investigate backup failure, take manual backup |
Each owner is responsible for executing the verification and initiating the recovery action if the expected result is not met. The checklist should be embedded in your monitoring system where possible (e.g., via cron jobs or Hop workflows).
Conclusion
Apache Hop offers a rich set of features for building robust data pipelines. The advanced concepts covered here - version control, environment separation, metadata injection, error handling, and recovery - are critical for production reliability. By following a disciplined approach of observing before changing, limiting blast radius, and verifying each step, you can avoid common pitfalls and keep your data flows running smoothly.
Start with a single low-risk action: run hop-conf --version and hop-conf --environment on your production server. Record the output. Then pick one pipeline and verify that it has proper error handling. Use the checklist to assign ownership and track improvements. Over time, these habits will build a solid operational foundation for your Hop deployments.