Translation and validation of the traditional Chinese Mobile App Rating Scale (MARS-TW) for evaluating mHealth apps
Highlight box
Key findings
• The traditional Chinese version of the Mobile App Rating Scale (MARS-TW) was developed and validated.
• MARS-TW demonstrated good reliability and validity, making it a suitable tool for evaluating traditional Chinese mHealth apps.
What is known and what is new?
• The MARS has been translated into multiple languages and has shown acceptable psychometric properties in prior studies.
• This study presents the first validated MARS-TW, adapted through a rigorous cross-cultural validation process.
What is the implication, and what should change now?
• MARS-TW provides a standardized evaluation tool for assessing traditional Chinese mHealth apps.
• Healthcare professionals and researchers can utilize MARS-TW to assess app quality, ensuring evidence-based selection and usage.
• Future research should explore its applicability in diverse health-related app categories beyond cancer care and hypertension management.
Introduction
Background
The World Health Organization states that using wireless technologies for public health and is an integral part of eHealth, which refers to the cost-effective and secure use of information and communication technologies in healthcare (1). Low cost, accessible mobile health applications (mHealth apps) have become endemic in supporting health and wellbeing. They offer accessible health information and services to facilitate health behaviors and prevent acute or chronic diseases (2,3).
There are more than one million mHealth apps in the world (4). In Taiwan, there is a growing number of mHealth apps related to diagnosis, treatment, data retrieving and decision making are categorized as medical software and regulated as medical devices (5). mHealth apps in Taiwan are not regulated. Users rely on App store scores and reviews in determining which apps to use. Unfortunately, mHealth apps reviews, provided by reliable originations are not commonly available (6,7).
Prior studies showed that few mHealth apps get validated for their quality and effectiveness. Bidargaddi, Musia (8) showed that of 539 oncology mHealth apps only 36.5% were validated. Kalke, Ginossar (9) also showed that amongst 302 breast cancer mHealth apps, less than one-third were content-validated. Haque and Rubya examined 193 mobile mental health applications and revealed that only 51.2% incorporated an evidence-based framework (10). Moreover, a scoping review of 29 digital health app studies indicated that 10 studies lacked empirical evidence, with app content frequently described as inappropriate, incorrect, or ambiguous, and the overall information quality often deemed unsuitable (11). Unvalidated mHealth apps may affect users’ health decisions and could cause serious health damage. Therefore, it is important to provide convenient and objective ways for general users or professionals to easily understand the quality of mHealth apps.
Mobile App Rating Scale (MARS)
The MARS, developed by Stoyanov et al. in 2015, is a validated tool for assessing mHealth app quality across multiple domains, including engagement, functionality, aesthetics, and information quality (12). The MARS includes a total of 29 items with three sections: objective quality ratings (19 items), subjective quality ratings (4 items) and perceived impact ratings (6 items). The scale also offers an app-classification matrix which can be used for the convenience of researchers to collect descriptive and technical information about the mHealth apps. Raters are asked to review the app description in the relevant apps stores to access this information.
The objective quality ratings section includes four subscales: “engagement” (5 items); “functionality” (4 items); “aesthetics” (3 items) and “information” (7 items). All items are rated on a five-point anchored Likert scale ranging from “1, inadequate” to “5, excellent”. The “information” domain also allows for an “N/A” option on five of its questions, for items which are inapplicable to apps which do not contain information (e.g., Games used for health). The mean score of the four subscales, excluding questions rated as “N/A”, measures the objective quality rating of the target mHealth apps.
The “subjective quality” section is used to determine the raters’ personal impression of the app. Lastly, the “perceived impact” section evaluate the rater-perceived effectiveness of the app to improve the knowledge, attitudes, intentions to change, as well as the likelihood of actual change in the target health behaviors.
The original MARS’ internal consistency Cronbach’s α was 0.90 and its inter-rater reliability intraclass correlation coefficient (ICC) was 0.79 (12,13). The scale has been used to determine the quality of a range of different mHealth apps (e.g., primary prevention, mindfulness, blood sugar control, medication compliance, bodyweight control, nutrition, sleep management, cancer management etc.) (13-16).
Purpose
The Chinese language is the second most popular language in the world (17) and it is categorized into Simplified and Traditional Chinese. Traditional Chinese is used as the official reading and writing system in parts of Chinese mainland, Taiwan, Hong Kong, Macao, Singapore, and Malaysia, etc. To our knowledge, there was no rating scale in Traditional Chinese for mHealth apps. Therefore, this study aimed to develop and validate a traditional Chinese of MARS (MARS-TW) by assessing its validity and reliability in evaluating traditional Chinese mHealth apps. The results would address the gap and provide mHealth apps developers and health care providers with a reference tool for developing and adopting safe and high quality mHealth apps.
Cultural equivalence of MARS translation and validation
The MARS has been translated into many languages including Italian, Spanish, Dutch, Arabic, German, French, Persian, Turkish and Japanese and validated through a variety of mHealth apps (e.g., primary prevention; physical activity; pregnancy; weight management). Translated versions have shown acceptable validity and reliability scores (e.g., Cronbach’s α 0.71–0.91; Intraclass correlations coefficients 0.67–0.96; Loevinger’s coefficient H 0.35–0.50) (14,18-25).
When transferring an instrument from its original development country to another, equivalence factors are important before adopting. Herdman, Fox-Rushby (26) provided six types of equivalence that should be considered, which includes: conceptual, item, semantic, operational, measurement and function (e.g., quality-of-life) to facilitate cross-cultural adoption, translation, validation, and reliability tests. The definitions and recommended methodologies for investigating each equivalence and were adopted in this study.
Methods
This study was authorized by the MARS authors on May 28, 2019 (12) to translate the original English version into the Traditional Chinese version, and then to validate its reliability and validity with local mHealth apps used in Taiwan since May, 2019. The multi-equivalence cultural adaptation model by Herdman, Fox-Rushby (26) and the development and validation of the different linguistic versions (e.g., Italian) were adopted into our three-phase study (I) conceptual and item equivalence; (II) semantic equivalence and (III) operational, measurement and functional equivalence.
Phase one: conceptual and item equivalence
In the first phase, three Taiwan experts, including medical informatics specialist, health care specialist and informatics technology specialist, were invited to review the English version of MARS. They used the four-point Likert scale to evaluate the relevance and acceptance of adopting MARS into the Taiwan culture by checking every item in each section independently (e.g., four points = highly acceptable; one point= not acceptable). The index of content validity (CVI) was used to evaluate the conceptual and item equivalence between MARS and MARS-TW; a CVI of more than 0.8 indicated that the three experts agreed and accepted the cross-culture translation in the next phase.
Phase two: semantic equivalence
In the second phase, the researcher (S.T.C.) forward-translated the initial MARS-TW, then two different Taiwan experts (informatic engineer and oncologist) were invited to review the initial MARS-TW independently. The four-point Likert scale was used to evaluate the acceptance of semantic contents of each item (e.g., four points = acceptable; three points = acceptable and need minor revision; two points = acceptable but need major revision; one point = not acceptable at all). The two experts also recommended providing alternative idioms and rephrases and then the focus meeting was held to achieve consensus on those translation problems that scored below three acceptance points. However, consensus was not reached for a few of the translated phrases. The research team decided to use the original phase as the keyword to review their semantic translation from references published in the Traditional Chinese journal in Taiwan and determine the final forward translation. Then the secondary revised MARS-TW was backward translated by professional bilingual translators. The backward-translated MARS was sent to the original developer (S.S.) for content evaluation and rechecked again by the research team to determine the final version of the MARS-TW.
Phase three: operational, measurement and functional equivalence
In the third phase, mHealth apps specifically for cancer care (CC) and hypertension self-management (HTM) with Chinese user interface were downloaded by two raters (e.g., researchers S.T.C. and Y.T.L.) to evaluate the validity and reliability of MARS-TW. There were three reasons for focusing on the CC and HTM mHealth apps. The first reason was considering the background experiences of the raters (e.g., S.T.C. was a graduate nurse student with oncology care experience and Y.T.L. was a community care graduate student) when reviewing these specific domain apps. The second reason was the samples size of any single target mHealth apps with Chinese interface language was less than the acceptable sample size (e.g., 41 apps) for validity and reliability tests (27). The third reason was to evaluate the generalizability of MARS-TW through a broader range of mHealth apps.
Retrieving target mHealth apps
To retrieve the target mHealth apps, on March 25 until March 30, 2021, the researcher (S.T.C.) used the App Annie (ATTN: data.ai Designated Agent, San Francisco, CA, USA) and Google web as the search platform with literature review-based keyword entries. With regard to the CC mHealth apps, keywords: cancer, cancer screen, tumor, cancer care, cancer self-care, cancer symptom, cancer related, pain, hospice care, hematology, and stem cell transplantation were used (28,29). The keywords used for HTM mHealth apps were hypertension, blood pressure, blood pressure self-management and blood pressure monitor (30,31).
The inclusion criteria of the target mHealth apps were: (I) used traditional Chinese interface language; (II) could be downloaded in Taiwan App stores (e.g., Android market; Apple store; Line store); (III) related to the target apps. The exclusion criteria of the App were: (I) apps designed for medical professionals and not for non-professional users; (II) medical device-related apps (e.g., retrieving data from blood sugar medical device or heart rate monitor). The 47 eligible apps were downloaded and both researchers installed them into their smart phones before MARS-TW rating.
Reliability test of the apps with MARS-TW
Before using MARS-TW, the two raters (S.T.C. and Y.T.L.) watched the free training video provided by the MARS developer (32), then 3 (7%) of all the eligible apps were randomly sampled for the pilot test. Each app was used at least 10 minutes and then rated with MARS-TW independently (12,14). After rating, the two raters checked the pilot results and consensus meetings were held for the non-consistent rating items. Cronbach’s α for internal consistency and ICC for the reliabilities of the two raters were tested. In this study, the Cronbach’s α levels were interpreted as excellent (α≥0.90), good (0.80≤α<0.90), acceptable (0.70≤α<0.80), doubtful (0.60≤α<0.70), bad (0.50≤α<0.60) and worst (α<0.50); the inter-rater reliability of items, subscales and MARS total scores were measured by means of ICCs using two-way mixed effects, average measures model with absolute agreement. ICC were interpreted as excellent (≥0.90), good (0.76–0.89), moderate (0.51–0.75) and poor (≤0.50) (12,33,34).
When the reliability test for the pilot apps was acceptable (e.g., Cronbach’s α ≥0.70 and ICC ≥0.50), the two raters then rated the remaining CC mHealth apps independently. A second consensus meeting was held to rating discrepancies between the two raters. Thereafter, the two raters then rated the remaining HTM mHealth apps independently.
Validity test of the apps with MARS-TW
Three validity tests were adopted to investigate the validity of MARS-TW. The first test was convergent validity which explored the conceptual similarity of each item in its own objective subscale. The test included item-subscale (I-S) correlation and item-total (I-T) scale of MARS-TW correlation. The Spearman’s ρ correlation was used for non-normally distributed ordinal and continuous variables and the coefficients >0.3 (I-S) and >0.2 (I-T) were regarded as satisfactory (14,34). The second test was divergent validity which explored the non-conceptual similarity between each objective subscale. A weak correlation (e.g., Pearson’s r coefficient) was expected and pairwise coefficients <0.7 were acceptable (14). The third test was concurrent validity which explored the rating between MARS-TW and star rating scores currently used. However, due to there is no alternative app quality rating scale available in Taiwan , the test was computed to check the correlation (e.g., Pearson’s r coefficients) among the star rating score (e.g., the 23rd item), quality mean score (e.g., mean score of the four objective subscales), App subjective quality (e.g., mean score of the MARS-TW subjective quality subscale) and rating systems available in app (e.g., star rating of the target app version in the App market and recorded in MARS-TW app classification) (12). Correlation coefficient were interpreted as very strong (between 0.90–1), strong (between 0.70–0.89), moderate (between 0.40–0.69), weak (between 0.10–0.39) and negligible (between 0.00–0.10) (35).
According to the results from the three phases, the research team would categorize the functional equivalence of MARS-TW into three broad types (e.g., complete match, acceptable match, no match) after comparing to the original MARS, and then justify and aggregate the results across cultures in Taiwan.
Results
Conceptual and item equivalence
Three Taiwanese experts independently evaluated the English version of MARS, yielding a content validity index (CVI) of 1.0, confirming its conceptual and item equivalence for use in Taiwan.
Semantic equivalence
The researcher translated MARS into the initial traditional Chinese version and then two experts reviewed it independently. The highest score of four points and the lowest score of three points were received from the two experts, and the average scores were 3.89 and 3.69, respectively. The initial results showed that the initial MARS-TW was acceptable and a total 37 words/phrases needed minor rephrasing (e.g., items with three points) according to the feedback of the two Taiwan experts. Then the researcher held the consensus meeting with the two experts and the inconsistent semantic translations were confirmed and the second version of MARS-TW was completed.
The second version of MARS-TW was back-translated by professional translators in Taiwan and reviewed by the MARS developer (S.S.). A total of six suggestions (see Table 1) were provided. The researcher and two translation review experts revised the second version of MARS-TW based on the suggestion of the MARS developer. However, there was still one word where consensus could not be reached for the translation (e.g., awareness); the researcher confirmed it by searching through traditional Chinese academic database within the last 5 years in Taiwan and the final version of MARS-TW was fully translated.
Table 1
| Section | Item | Backward translation | Suggestions |
|---|---|---|---|
| Engagement | 1 | Interesting/entertaining | Fun |
| 3 | Volume level | Sound | |
| Functionality | 8 | Smoothness | Flow logic |
| Aesthetics | 12 | Visual effects | Visual appeal |
| Information | 18 | Companies who stand to gain from this app | Companies who stand to gain from this app being misleading |
| Perceived impact | 1 | Cognitive | Awareness |
MARS-TW, traditional Chinese version of the Mobile App Rating Scale.
Operational, measurement and functional equivalence
Target mHealth apps
A total of 47 target mHealth apps (e.g., 25 CC apps and 22 HTM apps) were included (see Figure 1). Three CC apps were piloted and the reliability test showed the internal consistency was acceptable (e.g., Cronbach’s α=0.83) but the inter-rater reliability was poor [e.g., ICC =0.42 with 95% confidence interval (CI): 0.05–0.99]. Then the two raters reviewed the pilot test results and improved the alignment of app ratings with consensus discussion and then the inter-rater reliability was acceptable (ICC ≥0.50), and the remaining target apps (e.g., 22 CC apps and 22 HTM apps) could be rated independently.
Of the 44 target mHealth apps, 24 were in Android market (CC =9; HTM =15) and 16 apps were in Apple App store (CC =9; HTM =7); only four Apps were in Line store (CC =4; HTM =0). According to their affiliation, only 9 apps could be identified as developed by commercial companies (CC =4; HTM =5). Most of the CC apps focused on physical health (n=12; 54.5%) and providing information/education strategies (n=21; 95.4%). All the HTM apps, focused on physical health with monitoring/tracking strategies.
For all target apps, the skewness statistics showed that each objective subscale rated by each rater was almost symmetric. The mean scores were between 2.74 (engagement subscale scored by rater 2 on HTM apps) and 4.55 (perceived impact mean scored by rater 1 on CC apps). There was significant difference (P<0.001) between the two raters on scoring engagement subscale and perceived impact mean score subscale with large effect size (0.86 vs. 0.92).
Reliability test
The MARS-TW demonstrated good inter-rater reliability, with ICC values of 0.80 for functionality, 0.84 for information, and 0.86 for overall app quality. Moderate reliability was observed for engagement (ICC =0.68), aesthetics (ICC =0.70), subjective quality (ICC =0.72) and perceived impact (ICC =0.62). The individual item ICC varied between 0.22–0.82, with a mean of .59 [standard deviation (SD) =0.13] (see Table 2). The lowest ICC (0.22) was observed for item 5 (target group).
Table 2
| Subscale | Item | CC apps | HTM apps | All | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ICC | 95% CI | P | ICC | 95% CI | P | ICC | 95% CI | P | |||||||||||
| L | U | L | U | L | U | ||||||||||||||
| Engagement | 1 | 0.63 | 0.12 | 0.85 | 0.01 | 0.52 | −0.16 | 0.81 | 0.01 | 0.57 | 0.20 | 0.77 | <0.001 | ||||||
| 2 | 0.46 | −0.26 | 0.77 | 0.08 | 0.45 | −0.16 | 0.76 | 0.05 | 0.42 | −0.05 | 0.68 | 0.04 | |||||||
| 3 | 0.69 | 0.13 | 0.88 | <0.001 | 0.70 | 0.14 | 0.89 | <0.001 | 0.69 | 0.22 | 0.86 | <0.001 | |||||||
| 4 | 0.64 | 0.12 | 0.85 | 0.01 | 0.70 | −0.06 | 0.90 | <0.001 | 0.67 | 0.39 | 0.82 | <0.001 | |||||||
| 5 | 0.17 | −0.68 | 0.62 | 0.32 | 0.32 | −0.26 | 0.68 | 0.09 | 0.22 | −0.23 | 0.54 | 0.14 | |||||||
| All | 0.75 | 0.41 | 0.90 | <0.001 | 0.66 | −0.23 | 0.90 | <0.001 | 0.68 | 0.12 | 0.86 | <0.001 | |||||||
| Functionality | 6 | 0.26 | −0.37 | 0.65 | 0.18 | 0.56 | −0.08 | 0.82 | 0.04 | 0.49 | 0.10 | 0.72 | 0.01 | ||||||
| 7 | 0.61 | 0.11 | 0.84 | 0.01 | 0.66 | 0.21 | 0.86 | <0.001 | 0.69 | 0.41 | 0.83 | <0.001 | |||||||
| 8 | 0.63 | 0.12 | 0.84 | 0.01 | 0.68 | 0.26 | 0.87 | 0.01 | 0.64 | 0.33 | 0.80 | <0.001 | |||||||
| 9 | 0.48 | −0.29 | 0.79 | 0.08 | 0.51 | −0.16 | 0.80 | 0.05 | 0.59 | 0.24 | 0.78 | <0.001 | |||||||
| All | 0.73 | 0.36 | 0.89 | <0.001 | 0.78 | 0.49 | 0.91 | <0.001 | 0.80 | 0.62 | 0.89 | <0.001 | |||||||
| Aesthetics | 10 | 0.33 | −0.67 | 0.73 | 0.19 | 0.53 | −0.08 | 0.80 | 0.01 | 0.43 | 0.00 | 0.68 | 0.02 | ||||||
| 11 | 0.51 | −0.22 | 0.80 | 0.06 | 0.42 | −0.44 | 0.76 | 0.12 | 0.51 | 0.09 | 0.73 | 0.01 | |||||||
| 12 | 0.66 | 0.18 | 0.86 | 0.01 | 0.68 | 0.26 | 0.87 | <0.001 | 0.69 | 0.44 | 0.83 | <0.001 | |||||||
| All | 0.69 | 0.23 | 0.87 | <0.001 | 0.68 | 0.25 | 0.86 | <0.001 | 0.70 | 0.45 | 0.84 | <0.001 | |||||||
| Information | 13 | 0.72 | 0.34 | 0.88 | <0.001 | 0.63 | 0.16 | 0.85 | 0.01 | 0.82 | 0.66 | 0.90 | <0.001 | ||||||
| 14 | 0.87 | 0.70 | 0.95 | <0.001 | 0.28 | −0.75 | 0.70 | 0.24 | 0.70 | 0.44 | 0.84 | <0.001 | |||||||
| 15 | 0.57 | 0.00 | 0.82 | 0.03 | 0.48 | −0.25 | 0.80 | 0.08 | 0.54 | 0.14 | 0.75 | 0.01 | |||||||
| 16 | 0.76 | 0.33 | 0.91 | <0.001 | 0.67 | −3.04 | 0.96 | 0.15 | 0.74 | 0.41 | 0.88 | <0.001 | |||||||
| 17 | 0.78 | 0.44 | 0.91 | <0.001 | 0.00 | −0.47 | 0.46 | 0.50 | 0.42 | −0.08 | 0.70 | 0.04 | |||||||
| 18 | 0.50 | −0.12 | 0.79 | 0.05 | 0.81 | 0.55 | 0.92 | <0.001 | 0.80 | 0.63 | 0.89 | <0.001 | |||||||
| All† | 0.90 | 0.75 | 0.96 | <0.001 | 0.74 | 0.40 | 0.89 | <0.001 | 0.84 | 0.65 | 0.88 | <0.001 | |||||||
| Quality mean score | 0.87 | 0.69 | 0.95 | <0.001 | 0.83 | 0.44 | 0.94 | <0.001 | 0.86 | 0.74 | 0.93 | <0.001 | |||||||
| Subjective quality | 20 | 0.54 | −0.05 | 0.80 | 0.04 | 0.57 | −0.06 | 0.82 | 0.03 | 0.59 | 0.26 | 0.78 | <0.001 | ||||||
| 21 | 0.78 | 0.48 | 0.91 | <0.001 | 0.44 | −0.17 | 0.75 | 0.04 | 0.66 | 0.34 | 0.82 | <0.001 | |||||||
| 22 | 0.54 | −0.03 | 0.80 | 0.02 | 0.64 | 0.17 | 0.85 | 0.01 | 0.56 | 0.21 | 0.76 | <0.001 | |||||||
| 23 | 0.42 | −0.31 | 0.75 | 0.10 | 0.57 | 0.04 | 0.82 | 0.02 | 0.55 | 0.20 | 0.75 | <0.001 | |||||||
| All | 0.68 | 0.22 | 0.87 | <0.001 | 0.77 | 0.44 | 0.90 | <0.001 | 0.72 | 0.48 | 0.85 | <0.001 | |||||||
| Perceived impact | 1 | 0.65 | 0.14 | 0.85 | <0.001 | 0.43 | −0.18 | 0.79 | <0.001 | 0.58 | −0.11 | 0.82 | <0.001 | ||||||
| 2 | 0.75 | 0.41 | 0.89 | <0.001 | 0.68 | 0.20 | 0.87 | 0.01 | 0.81 | 0.66 | 0.90 | <0.001 | |||||||
| 3 | 0.57 | −0.14 | 0.83 | <0.001 | 0.42 | −0.24 | 0.77 | 0.01 | 0.57 | −0.17 | 0.82 | <0.001 | |||||||
| 4 | 0.55 | −0.18 | 0.83 | <0.001 | 0.36 | −0.23 | 0.71 | 0.06 | 0.53 | −0.10 | 0.79 | <0.001 | |||||||
| 5 | 0.59 | 0.04 | 0.83 | 0.01 | 0.64 | 0.17 | 0.85 | 0.01 | 0.59 | 0.24 | 0.78 | <0.001 | |||||||
| 6 | 0.60 | −0.19 | 0.86 | <0.001 | −0.17 | −1.83 | 0.52 | 0.64 | 0.52 | 0.08 | 0.75 | <0.001 | |||||||
| All | 0.58 | −0.21 | 0.85 | <0.001 | 0.61 | −0.18 | 0.86 | <0.001 | 0.62 | −0.11 | 0.85 | <0.001 | |||||||
†, item 19 was excluded from all calculations because of lack of ratings. CC, cancer care; CI, confidence interval; HTM, hypertension self-management; ICC, intraclass correlation coefficient; L, lower; U, upper.
The MARS-TW presented with excellent internal consistency levels (Cronbach’s α=0.84 with both raters). The subscale internal consistency scores ranged between 0.63 on the information and functionality subscales and 0.84 on the aesthetics subscale (see Table 3).
Table 3
| Subscale | Item | CC | HTM | All | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rater 1 | Rater 2 | Rater 1 | Rater 2 | Rater 1 | Rater 2 | |||||||||||||
| I-S | I-T | I-S | I-T | I-S | I-T | I-S | I-T | I-S | I-T | I-S | I-T | |||||||
| Engagement | 1 | 0.85 | 0.53 | 0.76 | 0.60 | 0.74 | 0.73 | 0.63 | 0.59 | 0.77 | 0.60 | 0.73 | 0.63 | |||||
| 2 | 0.70 | 0.37 | 0.74 | 0.52 | 0.64 | 0.64 | 0.62 | 0.76 | 0.66 | 0.46 | 0.72 | 0.71 | ||||||
| 3 | 0.32 | 0.28 | 0.60 | 0.49 | 0.87 | 0.62 | 0.81 | 0.59 | 0.71 | 0.44 | 0.68 | 0.53 | ||||||
| 4 | 0.85 | 0.77 | 0.81 | 0.62 | 0.88 | 0.69 | 0.72 | 0.43 | 0.87 | 0.68 | 0.79 | 0.58 | ||||||
| 5 | 0.67 | 0.65 | 0.45 | 0.56 | 0.75 | 0.75 | 0.35 | 0.39 | 0.65 | 0.64 | 0.39 | 0.43 | ||||||
| α | 0.65† | 0.69† | 0.79† | 0.64† | 0.75† | 0.69† | ||||||||||||
| Functionality | 6 | 0.72 | 0.50 | 0.76 | 0.34 | 0.37 | −0.14 | 0.52 | 0.34 | 0.63 | 0.29 | 0.75 | 0.48 | |||||
| 7 | 0.68 | 0.15 | 0.72 | 0.36 | 0.70 | 0.61 | 0.51 | 0.19 | 0.73 | 0.41 | 0.66 | 0.36 | ||||||
| 8 | 0.69 | 0.53 | 0.77 | 0.31 | 0.49 | 0.57 | 0.87 | 0.59 | 0.62 | 0.57 | 0.68 | 0.37 | ||||||
| 9 | 0.72 | 0.75 | 0.66 | 0.67 | 0.72 | 0.44 | 0.49 | 0.35 | 0.76 | 0.66 | 0.67 | 0.59 | ||||||
| α | 0.65† | 0.69† | 0.25† | 0.40† | 0.62† | 0.63† | ||||||||||||
| Aesthetics | 10 | 0.84 | 0.73 | 0.74 | 0.65 | 0.65 | 0.71 | 0.93 | 0.81 | 0.76 | 0.65 | 0.86 | 0.76 | |||||
| 11 | 0.93 | 0.80 | 0.91 | 0.74 | 0.90 | 0.45 | 0.94 | 0.86 | 0.91 | 0.67 | 0.93 | 0.83 | ||||||
| 12 | 0.90 | 0.76 | 0.89 | 0.76 | 0.62 | 0.26 | 0.75 | 0.71 | 0.83 | 0.61 | 0.82 | 0.73 | ||||||
| α | 0.87† | 0.80† | 0.57† | 0.84† | 0.78† | 0.84† | ||||||||||||
| Information‡ | 13 | 0.55 | 0.02* | 0.64 | 0.19* | 0.79 | 0.63 | 0.47 | 0.29* | 0.47 | 0.17* | 0.38 | −0.10* | |||||
| 14 | 0.71 | 0.57 | 0.85 | 0.61 | 0.60 | 0.47 | 0.61 | 0.57 | 0.61 | 0.47 | 0.79 | 0.56 | ||||||
| 15 | 0.83 | 0.67 | 0.75 | 0.62 | 0.64 | 0.60 | 0.64 | 0.25 | 0.75 | 0.65 | 0.71 | 0.48 | ||||||
| 16 | 0.77 | 0.61 | 0.80 | 0.69 | 0.65* | 0.08* | 0.39* | 0.64 | 0.71 | 0.53 | 0.70 | 0.66 | ||||||
| 17 | 0.75 | 0.61 | 0.73 | 0.48 | 0.68 | 0.48 | 0.41* | 0.51 | 0.68 | 0.58 | 0.66 | 0.40 | ||||||
| 18 | 0.16* | −0.03* | 0.54 | 0.41* | 0.38* | 0.18* | 0.54 | 0.44 | 0.31 | 0.17* | 0.44 | 0.54 | ||||||
| α | 0.72† | 0.81† | −1.84† | 0.37† | 0.63† | 0.63† | ||||||||||||
| Quality mean score† | α | 0.83† | 0.83† | 0.81† | 0.84† | 0.84† | 0.84† | |||||||||||
†, Cronbach’s α coefficients, by rater and subscale; ‡, item 19 was excluded from all calculations because of lack of ratings; *, P>0.05; CC, cancer care; HTM, hypertension self-management; I-S, item-subscale; I-T, item-total.
Validity test
The convergent validity tests showed that the correlation coefficient level of I-S was between 0.31 to 0.95 and I-T was 0.10 to 0.83 and most items were regarded as satisfactory (see Table 3).
The divergent validity test showed that the Pearson’s r correlation coefficients level of paired objective subscales was mostly <0.70 and were acceptable. The concurrent validity test showed that the correlation coefficient level was significant among the MARS-TW star rating, quality mean score and subjective quality mean score (r>0.52, P<0.05) (see Table 4). In line with scores presented by the original MARS study (12) the MARS-TW ratings did not show a significant correlation with the star rating scores in the App store. Three tests showed that the MARS-TW had acceptable convergent, divergent, and concurrent validity. According to the results from the three phases, MARS-TW was categorized as acceptable match across cultures in Taiwan after comparing it to the original MARS.
Table 4
| Item | Apps | MARS-TW star rating (N=23) | Quality mean score | Subjective quality mean score | Rating systems in App stores | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CC | HTM | All | CC | HTM | All | CC | HTM | All | CC† | HTM‡ | All | |||||
| MARS-TW star rating (n=23) | CC | – | 0.77** | 0.82** | 0.43 | |||||||||||
| HTM | – | 0.52* | 0.65** | −0.04 | ||||||||||||
| All | – | 0.62** | 0.74** | 0.45* | ||||||||||||
| Quality mean score | CC | 0.69* | – | 0.83** | 0.08 | |||||||||||
| HTM | 0.89** | – | 0.75** | 0.48 | ||||||||||||
| All | 0.83** | – | 0.71** | 0.36 | ||||||||||||
| Subjective quality mean score | CC | 0.89** | 0.66* | – | 0.40 | |||||||||||
| HTM | 0.97** | 0.89** | – | 0.31 | ||||||||||||
| All | 0.93** | 0.77** | – | 0.35 | ||||||||||||
| Rating systems in App stores | CC† | −0.34 | −0.07 | −0.04 | – | |||||||||||
| HTM‡ | 0.31 | 0.49* | 0.35 | – | ||||||||||||
| All | 0.06 | 0.32 | 0.06 | – | ||||||||||||
Rater 1: upper right triangle; Rater 2: lower left triangle. †, number of CC apps with rating systems available in App store =12; ‡, number of HTM apps with rating systems available in App store =17. *, P<0.05; **, P<0.01; CC, cancer care; HTM, hypertension self-management; MARS-TW, traditional Chinese version of the Mobile App Rating Scale.
Discussion
Major findings on MARS-TW culture equivalence
According to the results from phase one, MARS was totally recognized by Taiwan experts (CVI =1) and achieved conceptual and item equivalence easily. In phase two, the research team used forward and backward translations, expert validation, and literature review, etc., to confirm the acceptable terminologies of MARS-TW. The process took some time (e.g., 3 months) to achieve semantic equivalence between MARS and MARS-TW.
In phase three, we found that there were insufficient single-themed target mHealth apps in the traditional Chinese App market for effective statistical power. Meanwhile, considering the background knowledge of the raters, two different themes of target mHealth apps were tested for operational and measurement equivalence. According to Table 3, the reliability tests results showed that after rating all our target mHealth apps, MARS-TW had good inter-rater reliability (e.g., ICC ≥0.76) and internal reliability (e.g., Cronbach’s α ≥0.80). However, each subscale of the different themes of the target mHealth apps also showed moderately doubtful inter-rater reliability and internality reliability.
According to Tables 2,3, in all the target Apps, the convergent validity of each item in engagement, functionality, aesthetics and information subscales was satisfactory (e.g., P of I-S was >0.3 and I-T was >0.2); the divergent validity of paired subscale was acceptable (r<0.70, P<0.05); the concurrent validity showed that correlations between the MARS-TW star rating, quality mean score and subjective quality mean score were moderate or higher (r>0.62, P<0.05). However, accepted validity results for different themes of the target mHealth apps also existed.
To our knowledge, the study had consistent acceptable reliability and validity tests with the original MARS and other language versions developed by other countries (12-14,21,23-25). The non-acceptable reliability and validity tests may be caused by the retrieving target mHealth apps. In this study, we choose the CC and HTM mHealth apps as our target apps in consideration of our raters’ professional health training background. There was significance difference between the two raters on rating the engagement subscale (Wilcoxon Signed-Rank test was −3.75<0.00). We found that it would be difficult to rate the “entertainment” and “interest” items in the engagement subscale in consideration of the information needs of App users (e.g., cancer treatment education or blood pressure monitor). Therefore, the target apps may not satisfy the framework of MARS. In addition, when the pilot test was applied in the CC apps but not in the HTM apps, non-consensus between the two raters may also exist.
According to results in Table 3, the Cronbach’s α coefficients of the HTM apps were bad from the two raters (α=−1.84 and 0.37) in the “information” subscale. The results may cause the “N/A” items to affect the internal reliability. Regarding the evidence items of the “information” subscale in MARS-TW, they were all excluded in this study as our rater had difficulty finding every target app in the published scientific literature. Challenges also existed in the other four non-available items of the “information” subscale (e.g., goals; quality of information, quantity of information, visual information) when rating the mHealth apps due to the domain knowledge limitation of the rater. However, these items were important mHealth App quality indexes for health professionals and would be great challenges when using MARS-TW in the future.
Comparing the translation process with prior studies
Previous studies have validated MARS in multiple languages, including Italian, Spanish, Dutch, Arabic, German, French, Persian, Turkish, and Japanese (12,18-25). It has been widely applied to systematically evaluate mobile applications across diverse health and educational domains, such as smoking cessation, self-management of temporomandibular joint disorders, hemodialysis care, liver disease management, caries prevention in children and adolescents, and anatomical education (36-41). Our findings align with prior research, demonstrating acceptable reliability and validity metrics across linguistic adaptations. The MARS-TW translation process was similar with the Italian and Spanish versions of MARS (14,18). There were some differences between the culture equivalence evaluation of MARS-TW with other languages. For the conceptual and item equivalence, the Italian version adopted the review with IT literature (on usability in particular) and consultation with IT specialists; the Spanish version adopted the review process with three experts (IT specialists and health science specialists); however, the German and France versions of MARS had no item equivalence validation. For the semantic equivalence, almost every target language adopted forward and backward translations. After backward translations, most of them were confirmed by the original authors except for the German and French versions, which adopted the comprehensibility evaluation by local researchers and nonacademic (21,22). The Arabic version of MARS was forward and backward translated by professional translation teams and reviewed by professional members of Mid-Eastern and North-African government and academic institutions (20). These differences showed the flexibility of the semantic equivalence process when scales not developed locally were adopted.
Limitations
The limitation of this study was the few specifically themed Traditional Chinese mHealth apps that were available in Taiwan. Our study targeted two different mHealth apps and the statistical power was only 0.53 (n=44). The other limitation of the study was that only two raters were included in the study. Both limitations may affect the inferentiality of psychometrics when adopting MARS-TW.
Conclusions
MARS-TW is a culturally adapted, reliable, and valid tool for assessing Traditional Chinese mHealth apps. Its application can enhance app quality assessments and support evidence-based decision-making in digital health. Future studies should explore its utility across various health domains. Before using MARS-TW to evaluate the target mHealth apps, it is recommended to become familiar with the conceptual definitions for each of the subscales for higher quality evaluation; moreover, confirmation of psychometrics for scientific evaluation purposes is also required. mHealth app developers can also use the MARS-TW as the guideline for ensuring quality.
Acknowledgments
We appreciate Stoyan Stoyanov for allowing us to translate the original MARS Scale into traditional Chinese. The authors wish to thank all the experts who contributed to the translation and contextualization of the MARS, in particular: Dr. Po-Lung Chang, Dr. Ting-Ting Lee, Ms. Ming-Chuan Kuo, Mr. Chien-Huai Chuang, and Mr. Chih-Yu Chang.
Footnote
Data Sharing Statement: Available at https://mhealth.amegroups.com/article/view/10.21037/mhealth-25-10/dss
Peer Review File: Available at https://mhealth.amegroups.com/article/view/10.21037/mhealth-25-10/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://mhealth.amegroups.com/article/view/10.21037/mhealth-25-10/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- World Health Organization. WHO guideline: recommendations on digital interventions for health system strengthening. Geneva: World Health Organization; 2019. Report No.: Licence: CC BY-NC-SA 3.0 IGO.
- McKay FH, Cheng C, Wright A, et al. Evaluating mobile phone applications for health behaviour change: A systematic review. J Telemed Telecare 2018;24:22-30.
- World Health Assembly. mHealth: use of appropriate digital technologies for public health: report by the Director-General. Geneva: World Health Organization; 2018.
- Statista. Number of mHealth apps available in the Apple App Store from 1st quarter 2015 to 3rd quarter 2022. Available online: https://www.statista.com/statistics/779919/health-apps-available-google-play-worldwide/
- Taiwan Food and Drug Administration. Guidance for Medical Software Classification 2015 [updated 20229/15]. Available online: https://ibmi.taiwan-healthcare.org/zh/news_detail.php?REFDOCTYPID=&REFDOCID=0rilwwxuhor3rxg0
- Bondaronek P, Slee A, Hamilton FL, et al. Relationship between popularity and the likely efficacy: an observational study based on a random selection on top-ranked physical activity apps. BMJ Open 2019;9:e027536.
- Chen J, Cade JE, Allman-Farinelli M. The Most Popular Smartphone Apps for Weight Loss: A Quality Assessment. JMIR Mhealth Uhealth 2015;3:e104.
- Bidargaddi N, Musiat P, Winsall M, et al. Efficacy of a Web-Based Guided Recommendation Service for a Curated List of Readily Available Mental Health and Well-Being Mobile Apps for Young People: Randomized Controlled Trial. J Med Internet Res 2017;19:e141.
- Kalke K, Ginossar T, Bentley JM, et al. Use of Evidence-Based Best Practices and Behavior Change Techniques in Breast Cancer Apps: Systematic Analysis. JMIR Mhealth Uhealth 2020;8:e14082.
- Haque MR, Rubya S. For an app supposed to make its users feel better, it sure is a joke-an analysis of user reviews of mobile mental health applications. Proceedings of the ACM on Human-Computer Interaction 2022;6:1-29.
- Giebel GD, Speckemeier C, Abels C, et al. Problems and Barriers Related to the Use of Digital Health Applications: Scoping Review. J Med Internet Res 2023;25:e43808.
- Stoyanov SR, Hides L, Kavanagh DJ, et al. Mobile app rating scale: a new tool for assessing the quality of health mobile apps. JMIR Mhealth Uhealth 2015;3:e27.
- Terhorst Y, Philippi P, Sander LB, et al. Validation of the Mobile Application Rating Scale (MARS). PLoS One 2020;15:e0241480.
- Domnich A, Arata L, Amicizia D, et al. Development and validation of the Italian version of the Mobile Application Rating Scale and its generalisability to apps targeting primary prevention. BMC Med Inform Decis Mak 2016;16:83.
- Stec MA, Arbour MW, Hines HF. Client-Centered Mobile Health Care Applications: Using the Mobile Application Rating Scale Instrument for Evidence-Based Evaluation. J Midwifery Womens Health 2019;64:324-9.
- Azad-Khaneghah P, Neubauer N, Miguel Cruz A, et al. Mobile health app usability and quality rating scales: a systematic review. Disabil Rehabil Assist Technol 2021;16:712-21.
- Farooq O. 5 Most spoken languages in the world 2021 [updated August 12, 2021 Available online: https://www.insidermonkey.com/blog/5-most-spoken-languages-in-the-world-972042/
- Martin Payo R, Fernandez Álvarez MM, Blanco Díaz M, et al. Spanish adaptation and validation of the Mobile Application Rating Scale questionnaire. Int J Med Inform 2019;129:95-9.
- Tency I, Van Hecke A, Derycke J, et al. editors. Development and validation of the Dutch version of the Mobile Application Rating Scale (MARS): a pilot study on pregnancy apps. 3rd CARE4–International Scientific Nursing and Midwifery Congress 2019, Date: 2019/02/04-2019/02/06, Location: Leuven; 2019.
- Bardus M, Awada N, Ghandour LA, et al. The Arabic Version of the Mobile App Rating Scale: Development and Validation Study. JMIR Mhealth Uhealth 2020;8:e16956.
- Messner EM, Terhorst Y, Barke A, et al. The German Version of the Mobile App Rating Scale (MARS-G): Development and Validation Study. JMIR Mhealth Uhealth 2020;8:e14479.
- Saliasi I, Martinon P, Darlington E, et al. Promoting Health via mHealth Applications Using a French Version of the Mobile App Rating Scale: Adaptation and Validation Study. JMIR Mhealth Uhealth 2021;9:e30480.
- Barzegari S, Sharifi Kia A, Bardus M, et al. The Persian Version of the Mobile Application Rating Scale (MARS-Fa): Translation and Validation Study. JMIR Form Res 2022;6:e42225.
- Mendi O, Kiymac Sari M, Stoyanov S, et al. Development and validation of the Turkish version of the Mobile App Rating Scale - MARS-TR. Int J Med Inform 2022;166:104843.
- Yamamoto K, Ito M, Sakata M, et al. Japanese Version of the Mobile App Rating Scale (MARS): Development and Validation. JMIR Mhealth Uhealth 2022;10:e33725.
- Herdman M, Fox-Rushby J, Badia X. A model of equivalence in the cultural adaptation of HRQoL instruments: the universalist approach. Qual Life Res 1998;7:323-35.
- Zou GY. Sample size formulas for estimating intraclass correlation coefficients with precision and assurance. Stat Med 2012;31:3972-81.
- Böhme C, von Osthoff MB, Frey K, et al. Development of a Rating Tool for Mobile Cancer Apps: Information Analysis and Formal and Content-Related Evaluation of Selected Cancer Apps. J Cancer Educ 2019;34:105-10.
- Lu DJ, Girgis M, David JM, et al. Evaluation of Mobile Health Applications to Track Patient-Reported Outcomes for Oncology Patients: A Systematic Review. Adv Radiat Oncol 2021;6:100576.
- Alessa T, Hawley MS, Hock ES, et al. Smartphone Apps to Support Self-Management of Hypertension: Review and Content Analysis. JMIR Mhealth Uhealth 2019;7:e13645.
- Jamshidnezhad A, Kabootarizadeh L, Hoseini SM. The Effects of Smartphone Applications on Patients Self-care with Hypertension: A Systematic Review Study. Acta Inform Med 2019;27:263-7.
- Stoyanov S. MARS training video. 2016. (Available upon request from the developer).
- Portney LG, Watkins MP. Foundations of clinical research: applications to practice: Pearson/Prentice Hall Upper Saddle River, NJ; 2009.
- Raykov T, Marcoulides GA. Introduction to psychometric theory. Routledge; 2011.
- Schober P, Boer C, Schwarte LA. Correlation Coefficients: Appropriate Use and Interpretation. Anesth Analg 2018;126:1763-8.
- Hong Q, Wei S, Duoliken H, et al. Application of Behavior Change Techniques and Rated Quality of Smoking Cessation Apps in China: Content Analysis. JMIR Mhealth Uhealth 2025;13:e56296.
- Lim A, Merner B, Iyer S, et al. Evaluation of temporomandibular disorder self-management apps in Australia: a systematic review to inform clinical use. Int J Qual Health Care 2025;37:mzaf024.
- Kermani F, Mahmoodi M, Nasiri MR, et al. Quality review and content analysis of liver complications mobile apps in Iran: A statistical and machine learning approach. Int J Med Inform 2025;197:105842.
- Rivera García GE, Cervantes López MJ, Ramírez Vázquez JC, et al. Reviewing Mobile Apps for Teaching Human Anatomy: Search and Quality Evaluation Study. JMIR Med Educ 2025;11:e64550.
- Esmaeeli E, Khorashadizadeh MS, Rahmani M. Mobile Applications for Hemodialysis: Evaluation Using the Mobile App Rating Scale (MARS). Semin Dial 2025;38:102-10.
- Kang YM, Yeo AN, Lee SY. Expert Usability Evaluation of a Mobile Application for Systematic Caries Management in Children and Adolescents: A Cross-sectional Study. Int J Clin Pediatr Dent 2024;17:1370-6.
Cite this article as: Chen ST, Lu YT, Stoyanov S, Liu PC, Chang KJ, Hou IC. Translation and validation of the traditional Chinese Mobile App Rating Scale (MARS-TW) for evaluating mHealth apps. mHealth 2025;11:55.

