A graphic featuring the text "meshIQ®" in a modern font. The design emphasizes clean lines and a minimalist aesthetic, suitable for branding or logo use.

Moving Mainframe Data to Snowflake and AWS Through Apache Kafka®: Treehouse Software and meshIQ

Jenaya Devan October 5, 2026

Treehouse Dataflow Toolkit moves mainframe data from Db2, VSAM, and IMS through Apache Kafka® pipelines into Snowflake and AWS targets. meshIQ keeps the streaming layer visible and stable. Together they give data science teams continuously updated enterprise data for AI and ML.

by Joseph Brady, Director of Business Development, Treehouse Software, Inc.and Jenaya Devan, Channel Sales Director at meshIQ

Key takeaways

  • Treehouse Dataflow Toolkit (TDT) moves mainframe and non-mainframe data from Apache Kafka® pipelines into Snowflake and AWS targets, with both bulk load and change data capture (CDC).
  • meshIQ provides subscription support and a unified management console for Apache Kafka®, so the streaming layer in the middle stays visible, supported and stable.
  • Together, they give data science teams a continuously updated, history-preserving copy of enterprise data for analytics, machine learning and AI.

These customers also have a crucial need to tap into today’s advanced data analytics platforms, where an ever-expanding array of analytics, BI, ML, and AI tools are available to generate vital insights from their enterprise’s data.  Data science teams are eagerly awaiting the arrival of critical data from their enterprise’s data sources to supercharge their predictive analytics and generative AI frameworks.

The hard part is the path in between. Data has to be replicated off the mainframe, streamed reliably, and landed in a form analysts can actually query. Treehouse Software and meshIQ each handle a piece of that path. Treehouse gets the data from the source into analytics-ready targets, and meshIQ makes the Apache Kafka® layer in the middle straightforward to run.

What is Treehouse Dataflow Toolkit (TDT)?

Treehouse Dataflow Toolkit (TDT) is a cloud-native, fully automated, turnkey solution that moves data from Apache Kafka® streaming pipelines into analytics-friendly targets, including Snowflake, Amazon Redshift, Amazon Athena/S3, and Amazon S3 Express One Zone. It supports both bulk load and change data capture (CDC). Learn more about TDT for mainframe data sources.

TDT is built as a set of AWS Lambda-based microservices, so data transfers are highly available, auto-scaling and event-driven. It is also more than a connector. Providing an innovative and robust Lambda-based microservices infrastructure, TDT automatically generates every target resource the transfer needs, including schemas, history tables, current views, user views, stages and file formats. Without that automation, teams can spend months designing and building those structures by hand. TDT can also create optional archiving infrastructure and Apache Iceberg tables.

How does mainframe data get from the source?

The data flow has three stages:

  1. Replicate. A mainframe data replication tool, such as Rocket Data Replicate and Sync (RDRS) or CONNX, publishes bulk-load and CDC data from sources like Db2, VSAM, Adabas, IMS, IDMS and Datacom to Apache Kafka®.
  2. Stream. Apache Kafka® carries that data reliably and at scale, whether it runs on Amazon MSK, Confluent or a self-managed cluster.
  3. Load. TDT consumes the data from Apache Kafka® and loads it into the target’s delta tables, which keep the full history of the source data from the moment synchronization began. Current views are maintained alongside them, so teams get both the history needed for trend, predictive and prescriptive analytics and an up-to-date snapshot of the source.
Architecture diagram showing an AWS cloud pipeline that connects mainframe data sources to TDT AWS microservices via meshIQ, with tools for replication, monitoring, data processing, and analytics.
Figure 1: meshIQ provides Apache Kafka® pipeline management and control for data delivery to TDT and on to Snowflake and AWS-based targets.
Architecture diagram showing Treehouse TDT automatically creating data-transfer resources: AWS mainframe replication feeds meshIQ and TDT microservices, with Snowflake schemas, tables, stages, file formats, and views.
Figure 2: TDT automatically creates all target structures (schemas, history tables, current views, user views, stages and file formats) via bulk load and CDC.

Why does the Apache Kafka® layer need its own management?

Apache Kafka® is what lets this architecture scale, but it also becomes a critical system in its own right. If a broker falls behind, partitions become under-replicated or consumer lag builds up, the data analysts depend on arrives late or not at all. Running Apache Kafka® well takes clear visibility into the cluster and the ability to act on what you see.

The meshIQ advantage

Treehouse Software’s TDT solutions fully support data transfers from mainframe and non-mainframe data sources to the meshIQ platform, which provides unified visibility, control, and modernization of messaging, event streaming, and B2B transactional flows across all middleware vendors, all environments, and at enterprise scale. Acting as a real-time intelligence layer across your entire ecosystem, meshIQ connects, monitors, and analyzes every message, event, and business transaction with forensic-level insight, so you can run incredibly lean operations with confidence and control.  You can build your own Apache Kafka® management environment in just a few steps, monitor your real-time data flow, and use a range of features designed for enterprise environments, including:

  • One console for every cluster. Manage multiple Apache Kafka® clusters across Confluent, Amazon MSK, Strimzi and self-managed deployments, as well as distributions such as IBM Event Streams and Cloudera Apache Kafka®.
  • Real-time monitoring. Dashboards track message throughput, broker health, consumer lag, partition status and request latency.
  • Automated rebalancing. Partitions are redistributed as clusters grow, without disrupting operations.
  • Access control and auditing. Granular permissions, simplified ACL management and full audit logging.
  • Anomaly detection. Unusual behavior, such as unauthorized login attempts or abnormal data flow, is flagged in real time.
  • Commercial subscription support for open-source Apache Kafka®.

For a TDT pipeline, this means the team can watch consumer lag on the topics TDT reads, catch an unhealthy broker early, and keep data flowing to Snowflake and AWS on schedule.

Dark cluster-monitoring dashboard with three green online-health gauges, latency bars, storage metrics, and an amber alert; technical, status-focused tone.

What do Treehouse Software and meshIQ deliver together? A powerful combined solution….

Together, Treehouse Software and meshIQ provide a complete path from legacy source to analytics-ready target. TDT automates the data movement and the build of every target resource, while meshIQ keeps the Apache Kafka® pipeline in between visible, supported and stable. The target platform continuously accrues the most current source data alongside its full history, which is exactly what data scientists need for trend analysis, predictive analytics, ML and AI work.

By eliminating the expense, complexity, and manual nature of managing multiple middleware platforms, meshIQ helps enterprises dramatically cut OPEX, resolve incidents and disputes up to 70% faster, reduce manual reconciliation efforts, and accelerate service delivery. The platform also provides a safe, fast, and reliable path to modernize middleware and migrate to modern architectures such as Apache Kafka®, Apache ActiveMQ®, and cloud-native environments—without disruption.

Frequently asked questions

Which data sources does TDT support?

On the mainframe, TDT works with data from Db2, VSAM, Adabas, IMS, IDMS, Datacom and flat files, published to Kafka by a partner replication tool. For non-mainframe sources, the TDT-DIRECT AWS Plugin supports PostgreSQL, SQL Server, Oracle and Db2.

Which targets can TDT load into?

Snowflake, Amazon Redshift, Amazon Athena/S3, Amazon S3 Express One Zone and Amazon Aurora PostgreSQL.

Does TDT replace my mainframe replication tool?

No. TDT works alongside replication tools such as RDRS and CONNX, which publish data to Kafka. TDT takes over from there.

Does meshIQ work only with open-source Apache Kafka®?

No. The meshIQ Console for Apache Kafka manages open-source Kafka and commercial distributions, including Confluent, Amazon MSK, Strimzi, IBM Event Streams and Cloudera Kafka.

About Treehouse Software
Treehouse Software has specialized in enterprise mainframe data integration since 1982, helping organizations move complex legacy data into modern platforms without sacrificing speed, security or best-practice adherence. Through TDT, Treehouse supports data delivery from a wide range of mainframe sources, including Db2, VSAM, Adabas, IMS, IDMS, Datacom and flat files, into Snowflake and AWS. Treehouse welcomes the opportunity to discuss your environment and help identify the right approach for your data modernization goals. Contact us for more information or to schedule a demonstration.

TDT is ©Treehouse Software, Inc. All rights reserved.

About meshIQ 
meshIQ provides management, observability and subscription support for the messaging and streaming middleware enterprises run on, including Apache Kafka®, IBM MQ™, Apache ActiveMQ®, RabbitMQ®, Solace and TIBCO. The meshIQ Console for Apache Kafka® gives teams one place to monitor, manage and scale Apache Kafka® across distributions and environments, backed by commercial subscription support. Contact us for more information or to schedule a demonstration, or start a free trial of the meshIQ Console for Apache Kafka®.

Cookies preferences

✕

Others

Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.

Necessary

Necessary
Necessary cookies are absolutely essential for the website to function properly. These cookies ensure basic functionalities and security features of the website, anonymously.

Advertisement

Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.

Analytics

Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.

Functional

Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.

Performance

Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.