📌 Introduction:
Shanghai Zhijian Information Technology Co., Ltd. (“Zhijian”) was jointly founded in January 2013 by executives from well-known IT companies including Taobao and Acorn. Zhijian is committed to providing enterprises with integrated omnichannel membership management solutions, offering consulting services and technical support for end-to-end membership operations such as member acquisition, member care, member marketing, and member service. Today, Zhijian serves 200 million users across more than 120 retail and e-commerce customers.
Zhijian changed its original Lambda architecture approach and built its data platform on Singdata Lakehouse, effectively reducing data redundancy and lowering overall TCO by 40%. With Singdata Lakehouse’s fully managed, integrated lakehouse platform, Zhijian eliminated the burden of self-managed operations. Singdata now supports the entire data pipeline, including data integration, processing, real-time BI, task monitoring, cluster stability, and other data management and operations work. The data team can focus more on discovering and creating business value from data. In addition, based on the capabilities of Singdata Lakehouse, Zhijian uses multi-tenancy and secure data authorization services to safely and flexibly authorize customer access to data assets, expanding the application scope of data assets and unlocking data value.
Author: Yu Yan, Product Manager at Zhijian
Current Platform Status
Zhijian is committed to building an omnichannel, full-scenario, and full-touchpoint digital operations platform for enterprises. Because it creates value for customers through data, Zhijian places great importance on data platform construction, using fresher data and a more efficient platform to support customers’ business growth and help them discover and realize the value of data. In the early stage of data platform construction, the technical team explored building a self-managed real-time data warehouse platform to support the Shopping Guide Assistant and Operations Assistant data applications that Zhijian provides to customer systems.
In these data applications, requirements for data timeliness became increasingly high. Customers raised demanding expectations for both the freshness of data from the moment it is generated in business systems to the moment it can be queried in BI reports, and for elastic concurrent BI query capabilities. As the business developed, bottlenecks in the self-managed real-time architecture gradually emerged. For example, when a tenant’s membership base reached the hundred-million level, the self-managed data platform could not effectively guarantee query speed or data freshness, and it also suffered from many data hotspot issues. Zhijian therefore re-examined the data platform architecture and summarized the problems as follows:
1. Compute resource demand has peaks and valleys, while fixed configurations struggle to adapt to fluctuating compute needs
As business scale expanded, the data platform’s demand for compute resources and storage capacity grew accordingly. To meet peak-period SLAs under a self-managed data platform model, the self-built cluster size would need to be matched to peak demand. Whether to significantly increase cost for peak performance became a difficult internal trade-off between meeting requirements and controlling cost.
2. A large amount of hidden development and operations cost was consumed
Beyond the hard cost of purchasing resources, a self-managed data warehouse also required substantial labor investment.
Building a complete data pipeline requires more than a cloud data warehouse. It also requires data integration, ETL, scheduling, and other modules working together. This meant data engineers had to spend significant effort assembling other components. For monitoring and troubleshooting across the full data pipeline, data engineers often had to investigate step by step, resulting in long cycles and delayed business releases. In addition, to match business requirements, each resource specification adjustment required deployment and coordination across other components. Scaling up or down also required data redistribution. For this reason, Zhijian wanted a low-barrier, pay-as-you-go data platform with a WYSIWYG experience—one that not only optimized the data engine architecture, but also strengthened data development and management capabilities.
3. The data architecture was complex and assembled from many components, while multiple data copies across multiple pipelines created data consistency problems
To support brand-side data integration, processing, monitoring, operations, and maintenance, a self-managed cloud data warehouse approach required using ECS clusters to build and optimize technical components such as MongoDB, Flink, and open-source real-time data warehouses. This meant the operations team had to assemble real-time, offline, metadata, data storage, security permission management, scheduling, data integration, and many other components. They also had to select, operate, and maintain versions for more than ten open-source components to build the data platform system, making platform selection, construction, and operations extremely costly. For example, the existing business system required real-time data warehouse query capabilities, while the BI system connected to MongoDB. If MongoDB data changed but the data warehouse was not updated in time, the business system’s targeting results and BI query results could become inconsistent.
As the business scaled, Zhijian’s need for compute resources, data analysis, and data mining also grew. Based on these issues, the team planned to build a holistic solution to improve data warehouse development efficiency. The goal was to keep costs controllable while meeting the business side’s requirements for stronger data analysis and mining performance. The main selection considerations were:
- Expand the radius of data business capabilities, support rapid business expansion, and cover omnichannel business scenarios.
- Improve efficiency while keeping the architecture simple and costs elastic and controllable.
- Provide high-concurrency SLA guarantees so that business systems and BI reporting systems have better consistency.
Exploration and Comparison of the Target Data Platform Positioning
In the middle of the year, Zhijian held internal discussions about the positioning of the data warehouse platform. In addition to solving the challenges mentioned above, the team also focused on two long-term aspects:
- Integrated full-process capabilities. The platform should cover data collection, storage, processing, analysis, visualization, and the full data pipeline, freeing up data engineers’ time and improving efficiency. This would save time and cost while improving service response efficiency for the business side.
- Data analysis capability of the data warehouse. With rapid business growth and continuous data volume expansion, Zhijian’s early architecture mainly relied on transactional databases to meet daily business needs. However, as business complexity increased—especially in advanced analytical scenarios such as multi-table joins, large-scale aggregation and statistical computation, and fast elastic concurrency—traditional transactional databases gradually exposed limitations in performance, scalability, and functional richness. These limitations made it difficult to obtain timely and accurate data analysis and insights. Zhijian needed to introduce a more powerful analytical data warehouse to efficiently handle complex data analysis scenarios and meet the need for deeper data insight.
During this period, Zhijian learned that Singdata Lakehouse is a fully managed, integrated SaaS platform. Its product capabilities matched Zhijian’s requirements for high elasticity, low operations burden, simple architecture, and controllable cost. The team therefore compared a self-managed cloud data warehouse with Singdata’s fully managed integrated data platform, summarized as follows:

After summarizing comparisons across multiple scenarios, Zhijian saw that Singdata, as a fully managed data platform, was easy to use, simple to operate, sufficiently capable of supporting the business, and able to simplify architecture. It stood out in particular for its ability to quickly adapt to and support multiple data application scenarios. Zhijian also paid close attention to reducing total cost of ownership, and test results showed that performance would not be affected. Therefore, Zhijian decided to adopt Singdata Lakehouse.
Results and Value
After completing the migration, Zhijian used Singdata Lakehouse for a period of time and gained deeper practical experience. The team was able to see more clearly how it created value for Zhijian’s business, summarized in the following points:
1. A fully managed integrated data platform simplified and unified the architecture, essentially removing hidden development and operations burdens
Zhijian migrated the SaaS-side data pipeline to Singdata Lakehouse, simplifying capabilities such as data integration, task monitoring, and cluster stability. Data management became more convenient, eliminating the complexity and risk of manual configuration. The team was able to focus on innovation around data business value. Data engineers no longer needed to work across different development environments or components, significantly freeing up the technical team’s energy.
During the migration, Zhijian tested Singdata’s SQL syntax compatibility. More than 99% of SQL syntax and tasks required no adjustment. The few syntax issues encountered were resolved through communication with Singdata’s technical team. Zhijian recognized Singdata’s professionalism and service capabilities in data platforms.
After the upgrade, Zhijian’s big data architecture became very streamlined:

2. Zhijian benefited from an elastic SaaS platform, using storage-compute separation and elastic scaling to avoid idle resource waste and reduce cost
Singdata Lakehouse provides on-demand elastic resources, supporting multi-cluster isolation, automatic start and stop, and on-demand automatic elastic concurrency. This means Zhijian no longer needs to worry about compute capacity during peak concurrency periods. Compute resources scale automatically according to workload. Compute clusters can be started in seconds and elastically expanded severalfold, ensuring the needs of business analysis and BI queries. After peak periods, excess compute resources are released and no longer occupy resources continuously, keeping costs under control.
In the past, before important holidays, Zhijian had to prepare extensively for deployment and operations, including machine requests, service deployment, and data layout, in order to guarantee SLAs during business peaks. Now, the team only needs to set the specification of compute resources (VC), and Lakehouse can automatically scale compute resources according to peak-period business volume changes to guarantee SLAs.
The figure below shows the dynamic relationship between Zhijian’s running jobs and VC elastic resources:

3. The Singdata engine combines simplicity and flexibility, balancing performance and cost effectively
Cost and performance have always been key concerns for Zhijian. Since connecting business scenarios to Singdata Lakehouse, Zhijian found that compute cost, measured against the previous cloud data warehouse, decreased by 40%. The system supports both offline and real-time analytics, and compared with specialized real-time data warehouses in the industry, it still maintains a performance advantage.
Innovating a Secure Data Authorization Service Model to Unlock Data Value
Zhijian’s customers, especially large enterprises with data processing and analysis capabilities, want to combine and analyze the data on Zhijian’s platform with their own data. However, due to the limitations of the self-managed data warehouse architecture, Zhijian previously could not provide external data services. Now, based on Singdata Lakehouse’s multi-tenancy and secure data authorization management capabilities, Zhijian can provide external data services while storing only one copy of the data. This fundamentally ensures the consistency of underlying shared data. On this basis, processed data can be safely opened and authorized to customers through secure permission authorization, creating a new data service model.

- Secure data authorization and sharing allow open views or data tables to be shared directly with corresponding recipients through authorization. Recipients can flexibly choose to further process the data directly or transfer it into their enterprise data warehouse.
- Data recipients can access Zhijian’s latest data anytime and anywhere on demand for BI presentation, transfer, and further processing. With minimum cost and maximum efficiency, they can integrate the data with their existing enterprise data warehouses and quickly support upper-layer business applications.
Summary and Reflections
As a small-to-medium-sized digital membership marketing management company in a rapid expansion stage, Zhijian’s experience upgrading to a fully managed integrated data platform may offer useful reference for data-native enterprises in the same field. Zhijian’s reflections and summary are as follows:
1. Self-built and free does not mean low cost
A self-managed cloud data warehouse may appear to have a lower cost, but in reality it contains many hidden costs. For example, frequent resource upgrades and downgrades, as well as adaptation between components, require significant time and resources. Purchasing mature big data products instead of building internally, and letting a professional team handle data architecture and technical services, can be a strong choice. When selecting big data products, companies can prioritize products with usage-based billing models, whose cost advantages become more significant as customer scale grows.
2. Use fully managed services to release development resources and focus on data applications and customer value
For fast-growing Digital Native enterprises, maintaining business expansion is the company’s top priority. Fully managed big data products can release development resources and allow teams to focus on the business side and customer value. Of course, product selection should also fully consider whether the product has automated operations capabilities, complete operations inspection, and monitoring and alerting mechanisms, so that it can reduce team operations pressure and bring attention back to the business itself.
3. Stay open and continue exploring how to unlock data value
Data asset operations are accelerating the release of enterprise data value, which is currently a trend in data governance development. Data circulation and sharing require more secure and open data infrastructure. If companies choose a self-managed cloud approach, they need to consider cost and constraints. Choosing a secure and open platform product makes it easier for enterprises to explore innovative business scenarios.
4. Explore a new model for secure data authorization
Zhijian is a data-oriented enterprise and has always hoped to help customers fully realize the maximum value of their data. Customers have long wanted to combine and compute their own data with the data collected and accumulated on Zhijian’s platform, using a secure, compliant, simple, and convenient authorization method, and directly compute based on authorized data. This is a newly envisioned service model that requires support from a platform with innovative capabilities. Singdata Lakehouse provides the multi-tenant secure authorization capabilities Zhijian needs, and it is built as a multi-cloud data platform, matching Zhijian’s requirements for secure data sharing. Therefore, Zhijian is working with Singdata to explore a new model for secure data authorization.



