CUSTOMER STORY

Yishu Pharmacy: The Data Journey of a Leading Pharmaceutical Retailer in Guizhou

Yishu Pharmacy: The Data Journey of a Leading Pharmaceutical Retailer in Guizhou

Yishu Pharmacy, officially Guizhou Yishu Pharmaceutical Co., Ltd., was founded in 1999. After more than two decades of growth, it has become a leading pharmaceutical retailer in Guizhou Province. As an integrated pharmaceutical company covering retail, wholesale, online sales, and offline sales, Yishu Pharmacy operates more than 1,200 directly managed stores across over 80 districts and counties in Guizhou, serves more than 5 million members, and reaches 50 million customer visits annually.

As its business continued to expand, Yishu Pharmacy accumulated massive volumes of operational data. Its previous data architecture used a cloud-vendor big data platform as the primary data warehouse, carrying 140 TB of existing data assets. It also used that vendor’s cloud data middle platform services for data modeling, metric management, and asset governance. This combined “data foundation + data middle platform” architecture supported Yishu Pharmacy’s core data application scenarios for several years, including business analysis BI reports, member operations analysis, and store sales monitoring.

However, as business complexity increased and cost-efficiency pressure intensified, the need for an upgrade became clear.

The Expected Value of a Data Middle Platform: Different Roles, Different Needs

When Yishu Pharmacy first introduced the data middle platform, it had high expectations. Different business roles each brought their own goals, hoping the platform would enable a more data-driven business upgrade.

RoleCore NeedWork to Be Done
CDO/CTO/Data leaderBuild an enterprise data asset system and enable data-driven decision-makingQuickly and effectively integrate scattered data and establish unified standards
IT teamBuild stable and efficient data infrastructureSupport a technical architecture that can adapt flexibly to business changes
Business analystQuickly access data for decision supportObtain required data quickly and easily
Business decision-makerUse data to optimize operations and reduce costsSee business impact quickly within limited budget resources

From the perspective of data leaders, the core requirement was to build an enterprise-level data asset system and enable genuinely data-driven decision-making. Their challenge was how to quickly and effectively integrate data scattered across business systems, while establishing unified data standards and definitions. The data middle platform’s promise of “unified metric definitions and elimination of data silos” was exactly the solution they hoped for.

The IT team cared more about the stability and efficiency of data infrastructure. They needed a technical architecture that could support flexible business changes, ensure stable output of core reports, and respond quickly to new requirements. In theory, the standardized development process and model management capabilities provided by the data middle platform could help them standardize development and improve collaboration.

Business analysts had the most direct requirement: get data quickly to support business decisions. They wanted to reduce their dependence on the IT team and access the business data they needed independently and conveniently. The “business self-service analytics” promoted by the data middle platform was a key selling point for them.

For business decision-makers, the focus was return on investment. They wanted to optimize operations and reduce operating costs through data, and they expected to see business results quickly within a limited budget. As a significant investment, the data middle platform had to prove its business value.

The Gap Between Ideal and Reality: Pain Points of the Data Middle Platform

After a period of practice, however, Yishu Pharmacy found a clear gap between the actual results of the data middle platform and its original expectations.

5.png

The most prominent problem was the paradox of development efficiency. The data middle platform emphasized standardized modeling and required all data development work to follow a standard path: subject-area division, conceptual model, logical model, physical model, and metric definition. This process was valuable for building core data assets, but the problem was that many business scenarios involved temporary, exploratory ad-hoc needs. A simple data query might only mean that a business user wanted to quickly check the distribution of data along one dimension, yet it still had to go through the full modeling process. This “overkill” approach made response speed slower than the traditional SQL script model, frustrating business users.

A deeper issue was the lack of perceived value. Objectively, the data middle platform did help Yishu Pharmacy build a relatively complete metric system and asset map, which was valuable for the data team’s internal asset management. But the real pain point was that frontline business users, including operations specialists, sales supervisors, and store managers, never logged into the data middle platform. They usually only looked at the final numbers and charts presented in BI reports, with no awareness of how data was processed, governed, or standardized. As a result, the data middle platform gradually became a tool used mainly by the data team. The data team spent substantial effort maintaining metric systems and optimizing model structures, but what business teams perceived was simply that “waiting time after submitting a request became longer.” They could not feel the value of the “middle platform” at all.

At the same time, cost pressure continued to accumulate. Computing costs rose linearly as data volumes grew, keeping the overall cost of the data platform high. As cost reduction and efficiency improvement became a company-wide priority, this data architecture, whose return on investment was difficult to quantify, became increasingly hard to justify for ongoing budget support.

The Core Motivation for the Upgrade: Aligning Cost with Business Value

Based on these pain points, Yishu Pharmacy decided to upgrade its data architecture. The essence of this upgrade was not simply replacing one technical component with another, but using a lighter and more flexible solution to replace an overly heavy system with unclear ROI.

Cost optimization was the key driver behind this architectural upgrade. Yishu Pharmacy’s goal was clear: make the cost of the data platform align with the business value it created. The previous model of “first building a large and comprehensive platform, then gradually discovering value” was no longer suitable. The company needed a new model built on pay-as-you-go usage and visible, measurable value.

Against this backdrop, Yishu Pharmacy began evaluating alternative solutions in the market and ultimately selected Singdata’s unified lakehouse platform.

Singdata Unified Lakehouse: The Design Philosophy of the New Architecture

Singdata’s unified lakehouse platform turns complex data infrastructure into a cloud-based online service delivered as SaaS.

First is the idea of unification. In traditional architectures, data lakes, data warehouses, and data middle platforms are often separate product components developed by different teams, deployed independently, and billed separately. The problems with this assembled architecture are obvious: data has to move between multiple components and be stored redundantly, interfaces between components must be continuously maintained, and any issue in one link can affect the entire data pipeline. Singdata Lakehouse integrates storage, compute, governance, and serving capabilities into one unified platform, removing component boundaries and fundamentally avoiding data redundancy and interface complexity.

Second is a transparent cost model. Singdata uses transparent billing based on actual resource consumption, allowing users to clearly see the cost generated by each compute task and each stored dataset. This online, cloud-native architecture with separated storage and compute helps data teams connect costs to specific business scenarios, so they can make resource allocation decisions that better match business needs.

Third is development flexibility. Singdata Lakehouse supports SQL as the unified development language across all scenarios. Offline batch processing, real-time stream processing, and interactive ad-hoc queries can all be implemented with the same SQL code. This means data developers no longer need to learn and maintain different technology stacks for different freshness scenarios, significantly reducing development and operations complexity.

In addition, Singdata provides services in a SaaS model, so users do not need to manage underlying infrastructure. The platform provides fully managed service assurance, freeing Yishu Pharmacy’s data team from heavy platform operations work and allowing them to focus on enabling the business with data.

Technical Explanation: What Is a Lakehouse Platform?

To understand the lakehouse platform, it helps to first look at the core idea behind this architectural paradigm.

Lakehouse is one of the most important technical evolutions in data architecture in recent years. It combines the flexibility of a Data Lake with the high performance of a Data Warehouse, aiming to deliver low-cost storage and efficient analytics within one unified architecture.

Traditional data lake architectures are built on low-cost object storage and can hold raw data in any format, but they lack the ACID transaction support, schema management, and query optimization capabilities of data warehouses. Traditional data warehouses provide strong analytical performance, but their closed storage formats and high storage costs make them less suitable for massive raw datasets.

The Lakehouse architecture has now become an industry consensus. On top of open data lake storage, next-generation table formats such as Apache Iceberg, which Singdata supports, introduce enterprise-grade capabilities including transaction support, schema evolution, and time travel. At the same time, intelligent caching, vectorized execution, adaptive optimization, and other technologies deliver query performance comparable to or even better than traditional data warehouses.

Singdata Lakehouse is built on this technical path and further integrates data integration, metadata management, data governance, and data serving capabilities, forming an out-of-the-box unified data platform.

Core Advantages of the Singdata Lakehouse Unified Architecture

Compared with a traditional Lambda architecture

1.png

Singdata Lakehouse supports unified Kappa architecture data warehouse construction

2.png

Quantifying the Cost Advantage

Cost optimization was one of the core considerations behind Yishu Pharmacy’s choice of Singdata Lakehouse. Comparing the cost structure of the original architecture with the new one makes the source of this advantage clearer.

The cost of the original “MC + data middle platform” architecture mainly consisted of storage fees billed by storage volume, MC compute fees billed by CU hours or usage, annual subscription fees for the data middle platform product, and redundant storage and compute overhead caused by data movement across different components. In addition, the operations labor required to maintain multiple components was also a hidden cost.

TCO = hardware + software + development + operations + governance

Rule of thumb: TCO is often at least 3x hardware cost

The cost structure of Singdata Unified Lakehouse is much simpler: storage fees based on actual consumption, compute fees based on actual consumption, and the platform SaaS service fee. Because the unified architecture eliminates data redundancy between components, and because intelligent storage tiering and compute optimization reduce processing cost per unit of data, the cost of processing data drops significantly.

Based on Yishu Pharmacy’s actual migration results, after the architecture was switched over with no disruption perceived by business users, the overall usage cost of the data platform fell by more than 50%. This result fully validated the value of the architectural upgrade.

Performance

Beyond cost advantages, data analytics performance was also an important consideration for Yishu Pharmacy during technology selection. As a pharmaceutical retail company, Yishu Pharmacy’s BI reports directly serve management and frontline business users, so query response speed is critical.

According to Singdata’s published TPC-DS performance test report, Singdata Lakehouse demonstrated industry-leading analytical performance in standard benchmark tests. In a 10 TB TPC-DS test, Singdata Lakehouse ranked among the top domestic lakehouse products in overall query performance. This performance comes from Singdata’s continuous investment in core technologies such as the query optimizer, vectorized execution engine, and intelligent caching.

3.png

For Yishu Pharmacy, this means that after migrating to Singdata Lakehouse, costs were optimized while BI report loading speed remained reliable, ensuring that the business user experience was not affected.

Seamless Cutover: Migration from MaxCompute to Singdata

Another key challenge in the architectural upgrade was achieving a smooth transition. Yishu Pharmacy’s BI reports directly serve executives and frontline business users, so any service interruption or data inconsistency could have had serious business impact.

Throughout the migration, supported by the full integrated capabilities of Singdata Lakehouse, Yishu Pharmacy’s data team completed a seamless self-service migration of data jobs to Singdata. Core data jobs and BI reports were fully switched to the new platform. More importantly, this migration was not only a technical architecture upgrade, but also an opportunity for data governance. Historical jobs without clear owners and redundant data were reviewed and cleaned up during the migration, making data assets leaner and more efficient.

Note: The Singdata engine can be inserted seamlessly and non-intrusively into existing open-source components, serving as a higher-performance engine for data analytics tasks.

4.png

From Data Middle Platform to Data Intelligence Infrastructure

Yishu Pharmacy’s architectural upgrade reflects an important shift in how enterprises build data platforms.

In Gartner’s 2024 Hype Cycle for Data, Analytics and AI in China, “data middle platform” was marked as a technology concept that is about to exit the historical stage. It is being replaced by the more pragmatic idea of “Data Infrastructure.” Data Infrastructure emphasizes foundational technical capabilities, including analytical databases, data integration, metadata management, data quality, and data services, as a reusable foundation for data analytics and AI applications, rather than a large and comprehensive application system.

Yishu Pharmacy’s practice confirms this trend. What enterprises truly need is not a feature-heavy “middle platform system” disconnected from business value, but a data infrastructure layer that is cost-controllable, performance-reliable, and flexible to use. Singdata Unified Lakehouse is a strong practice of this philosophy.

Looking ahead, Yishu Pharmacy will further explore the integration of data and AI on top of Singdata Lakehouse. In addition to foundational capabilities such as lakehouse and OLAP, Singdata Lakehouse also includes vector database and scalar search capabilities, and provides DataGPT and other data intelligence application capabilities built with AIGC technology, giving enterprises a solid foundation for data innovation in the AI era.

Farewell to fragmentation, return to the essentials: aligning data cost with business value is the core outcome of Yishu Pharmacy’s architectural upgrade, and it is also the value proposition of Singdata’s unified lakehouse platform.