{"id":75636,"date":"2020-03-27T07:42:42","date_gmt":"2020-03-27T05:42:42","guid":{"rendered":"https:\/\/prohoster.info\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki"},"modified":"2020-03-27T07:42:42","modified_gmt":"2020-03-27T05:42:42","slug":"strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki","status":"publish","type":"post","link":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki","title":{"rendered":"Structuring Risks and Solutions When Using Big Data for Obtaining Official Statistics","gt_translate_keys":[{"key":"rendered","format":"text"}]},"content":{"rendered":"<p><i>Translator's Foreword<\/i><\/p>\n<p>I was intrigued by the material, primarily because of the table below:<\/p>\n<p><img decoding=\"async\" alt=\"Structuring Risks and Solutions When Using Big Data for Obtaining Official Statistics\" src=\"\/wp-content\/uploads\/2020\/03\/e39f7d729696e704d437fca4d74b3d9d.jpeg\" style=\"display:block;margin: 0 auto;\" \/><br \/>\n<br \/>\nConsidering that statisticians (and Russians, at a genetic level) do not particularly favor anything that deviates from linear relationships, these individuals managed to advocate for the use of an activation function in a parabolic form to assess the risk associated with using Big Data in official statistics. Well done. Naturally, statisticians have added their note to this work \u2013 '1 Any errors and omissions are the sole responsibility of the authors. The opinions expressed in this document are personal and do not necessarily reflect the official position of the European Commission.' However, the work was published. I believe that is sufficient for today, and they (the authors) did not prohibit anyone from finding their scales in these aspects.<\/p>\n<p>The work is structured enough to distinguish where and how statistical methods differ from research methods for Big Data. In my opinion, the greatest benefit of this work will be in discussions with clients and for countering their statements like:<\/p>\n<p> \u2014 We collect statistics ourselves; what more do you wish to investigate?<br \/>\n \u2014 Present your results in such a way that we can reconcile them with our statistics. In this matter, the authors suggest that it would be worthwhile to read this work (3 <noindex><a rel=\"nofollow\" href=\"http:\/\/www1.unece.org\/stat\/platform\/download\/attachments\/99484307\/Virtual%20Sprint%20Big%20Data%20paper.docx?version=1&amp;modificationDate=1395217470975&amp;api=v2\">How big is Big Data? Exploring the role of Big Data in Official Statistics<\/a><\/noindex>)<\/p>\n<p>In this work, the authors express their view on the level of risk. This parameter is placed in parentheses, not to be confused with a reference to sources.<\/p>\n<p>Second observation. The authors use the term BDS \u2013 this is analogous to the concept of Big Data. (presumably a nod to official statistics).<br \/>\n<noindex><a rel=\"nofollow\" name=\"habracut\"><\/a><\/noindex><br \/>\n<i>Authors' Foreword<\/i><\/p>\n<p>An increasing number of statistical agencies are exploring the possibility of using large data sources for the preparation of official statistics. Currently, there are only a few examples where these sources have been fully integrated into actual statistical production. Consequently, the full extent of the implications brought about by their integration remains unclear. In the meantime, initial attempts have been made to analyze the conditions and impact of big data on various aspects of statistical production, such as quality or methodology. Recently, a target group has developed a qualitative framework for the preparation of statistical data based on big data in the context of the Economic Commission for Europe\u2019s big data project of the United Nations (UNECE). According to the European Statistical Code of Practice, providing high-quality statistical information is a primary objective of statistical agencies. Given that risk is defined as the impact of uncertainty on objectives (for example, by the International Organization for Standardization ISO 31000), we found it appropriate to categorize risks according to the dimensions of quality they affect. <br \/>\nThe proposed quality framework for statistical data derived from large data sources offers a structured representation of quality associated with all stages of the statistical business process, and thus can serve as the basis for comprehensive assessment and management of risks related to these new data sources. It introduces new quality dimensions that are specific to or highly relevant when using big data for official statistics, such as the institutional\/business environment or complexity. By employing these new quality dimensions, risks associated with the use of large data sources in official statistics can be identified more systematically.<\/p>\n<p>In this work, we aim to identify the risks posed by the use of big data in the context of official statistics. We adopt a systematic approach to defining risks within the proposed quality framework. Focusing on newly proposed quality metrics, we can describe risks that are currently absent or not impacting the production of official statistics. At the same time, we can identify existing risks that will be assessed quite differently when using big data for statistical purposes. We then move further in the risk management cycle and provide an assessment of the likelihood and impact of these risks. Since risk assessment entails subjectivity in attributing probability and impact to various risks, we measure the agreement among dozens of different stakeholders provided independently. We then offer risk mitigation options in accordance with four main categories: avoidance, reduction, sharing, and retention. According to ISO, one of the principles of risk management should be value creation, meaning that the resources for risk reduction should be less than those for inaction. In line with this principle, we will finally evaluate the potential impact of some risk mitigation measures on the quality of the final outcomes, aiming for a more comprehensive assessment of the use of Big Data for official statistics.<\/p>\n<h2>1. Introduction<\/h2>\n<p><\/p>\n<h4>1.1. Background<\/h4>\n<p>\nThe development of big data was characterized by Kenneth Neil Cukier and Viktor Mayer-Sch\u00f6nberger in their article \"The Growth of Big Data\" (2. <noindex><a rel=\"nofollow\" href=\"http:\/\/www.foreignaffairs.com\/articles\/139104\/kenneth-neil-cukier-and-viktor-mayer-schoenberger\/therise-of-big-data\">www.foreignaffairs.com\/articles\/139104\/kenneth-neil-cukier-and-viktor-mayer-schoenberger\/the-rise-of-big-data<\/a><\/noindex>) the term \"datafication.\" Datafication is described as the process of taking all aspects of life and turning them into data. For example, Facebook provides personal networks, sensors for all kinds of environmental conditions, smartphones for personal communication and movements, and wearable data for personal conditions. This leads to almost ubiquitous data collection and availability.<\/p>\n<p>As in many other sectors, official statistics have only recently begun to discuss the issue of big data at a strategic level. There is still no common and widely accepted understanding of the way forward, whether it is a challenge or an opportunity, whether it is small or large, etc. Within the High-Level Group on Modernization of Statistical Production and Services (3 How big is Big Data? Exploring the role of Big Data in Official Statistics: <noindex><a rel=\"nofollow\" href=\"http:\/\/www1.unece.org\/stat\/platform\/download\/attachments\/99484307\/Virtual%20Sprint%20Big%20Data%20paper.docx?version=1&amp;modificationDate=1395217470975&amp;api=v2\">www1.unece.org\/stat\/platform\/download\/attachments\/99484307\/VirtualSprintBigDatapaper.docx?version=1&amp;modificationDate=1395217470975&amp;api=v2<\/a><\/noindex>), a First SWOT analysis, accompanied by a rough risk\/benefit analysis, was conducted. It was noted that \"a full risk analysis would also include aspects such as likelihood and impact, and may also be expanded to identify risk mitigation and management strategies.\"<\/p>\n<p>Although this document is far from a complete risk analysis, it aims to improve the situation precisely by creating the first structured overview. We would like to emphasize that this overview should be seen as a starting point for stimulating general discussion within the Official Statistical Community (OSC).<\/p>\n<h4>1.2. Scope<\/h4>\n<p>\nThis article is dedicated solely to risks, excluding not only benefits but also strengths and weaknesses, opportunities, and threats. This means that \"risks of inaction\" (for example, the risk that the OSC will become uncompetitive with other players if it is not modernized) do not fall within the scope; these are rather threats. Instead, we try to highlight the risks that may arise (a) if the OSC takes advantage of the opportunities provided by big data and begins to develop or improve a specific \"Big Data-based Official Statistics Product\" (BOSP); (b) risks to the new \"business as usual,\" namely risks for official statistics based on \"big data\" production. (Since all production of official statistics involves risks, we limit ourselves to (b) risks specific to \"big data risks,\" i.e., risks that do not exist or are minimal for the \"traditional\" process of collecting official statistics.)<\/p>\n<h4>1.3. Structure<\/h4>\n<p>\nIn Section 2, we present the fundamental principles related to this task, starting with the clearly necessary foundation for risk management and risk governance (Section 2.1). We also introduce a preliminary quality structure for statistical data derived from big data (Section 2.2), as linking the quality structure with risks serves two purposes:<\/p>\n<ul>\n<li>It sets the context for defining risks. Certain quality indicators, along with the considered characteristics, express the values of the entity that are deemed important and critical for delivering services to clients and users.<\/li>\n<li>This allows for the assignment of specific risks to quality measurements that are embedded in overall hyperspaces and linked to particular stages in the production process of statistical products.<\/li>\n<\/ul>\n<p>\nIn Sections 3, 4, 5, and 6, we present the risks identified thus far in various contexts (The business case documents of the ESS Big Data project as well as in ESSnets contain a list of risks partially related to the project and partially to using big data sources for statistical purposes. The document 'A suggested Framework for the Quality of Big Data' mentions some risks related to quality dimensions.). Here, we use a classification of data access, legal environment, privacy and data security, as well as skills; reorganization according to the quality structure of statistics obtained from big data (Section 2.2) should be considered as soon as this structure reaches a more finalized status. For each identified risk, we (i) provide an assessment of probability and impact (as per Section 2.1.3) and (ii) propose risk mitigation and management strategies (see Section 2.1.4).<\/p>\n<p>In the end, we will discuss our conclusions and outline some next steps in Section 7.<\/p>\n<h2>2. Fundamentals<\/h2>\n<p><\/p>\n<h4>2.1. Risks and Risk Management<\/h4>\n<p>\nAccording to ISO 31000: 20095, risk is defined as \"the effect of uncertainty on objectives.\" This means that objectives must be established or known before risks can be identified. These objectives are typically defined in the context of the relevant organization\u2019s institutional environment. An additional important consideration is that risks carry a characteristic of uncertainty, meaning it is unclear whether the described event will occur. Thus, risks are measured in terms of the likelihood of an event occurring and its consequences, i.e., the impact that this event has on achieving the objectives. Risk assessment should provide more objective information that ultimately allows for a proper balance between realizing profit opportunities and minimizing adverse impacts. Risk management is an integral part of management practice and an important element of good corporate governance (6 Statistics Canada: 2014-2015 report on Plans and Priorities, <noindex><a rel=\"nofollow\" href=\"http:\/\/www.statcan.gc.ca\/aboutapercu\/rpp\/2014-2015\/s01p06-eng.htm\">www.statcan.gc.ca\/aboutapercu\/rpp\/2014-2015\/s01p06-eng.htm<\/a><\/noindex>). This is an iterative process that ideally allows for continuous improvement of decision-making processes and fosters ongoing performance enhancement.<\/p>\n<p>Risks are also linked to quality. The implementation of a quality system should enable the leveraging of opportunities provided by various sources and methodologies to achieve an outcome of a certain level of quality in the sense that this outcome meets user needs. Like risks, quality levels can be derived from the institutional environment and the objectives of specific organizations. In this context, the institutional environment defines the overall level of risk that the organization is willing to undertake to achieve its goals.<\/p>\n<p>The process of risk assessment and management can be broken down into various stages that include establishing the context, identifying risks, analyzing risks in terms of likelihood and impact, assessing risks, and finally, risk treatment.<\/p>\n<h4>2.1.1. Institutional Context<\/h4>\n<p>\nThe first step is to establish a strategic, organizational context and risk management framework within which the remainder of the process will take place. This includes setting the criteria by which risks will be evaluated and defining the analysis structure.<\/p>\n<h4>2.1.2. Risk Identification<\/h4>\n<p>\nIn the second stage, events that may impact the achievement of set objectives must be identified. Identification should include questions related to the type of risks, the timing of the events, locations, and how these events may prevent, worsen, delay, or enhance the attainment of goals.<\/p>\n<h4>2.1.3. Risk Assessment<\/h4>\n<p>\nThe next step involves identifying existing controls and analyzing risks in terms of likelihood and potential consequences. In the context of this article, the probability of risks is rated on a scale from 1 (unlikely) to 5 (frequent). The impact of events is measured on a scale from 1 (insignificant) to 5 (extreme). As shown in Table 1, the product of probability and impact yields a \"risk level\" ranging from 1 to 25.<\/p>\n<p><img decoding=\"async\" alt=\"Structuring Risks and Solutions When Using Big Data for Obtaining Official Statistics\" src=\"\/wp-content\/uploads\/2020\/03\/486877e0bd16137b74227392fe85957c.jpeg\" style=\"display:block;margin: 0 auto;\" \/><br \/>\n<br \/>\nEvaluated risk levels can be compared with pre-defined criteria to establish a balance between potential benefits and adverse outcomes. This allows for judgments about management priorities.<\/p>\n<p><img decoding=\"async\" alt=\"Structuring Risks and Solutions When Using Big Data for Obtaining Official Statistics\" src=\"\/wp-content\/uploads\/2020\/03\/b49cbaff80da9d93fdfedf7a94750b78.jpeg\" style=\"display:block;margin: 0 auto;\" \/><br \/>\n<br \/>\nPriority for action should be given to critical risks (see Table 2), namely those that may occur and have serious or extreme consequences for the organization's objectives.<\/p>\n<h4>2.1.4. Risk Response<\/h4>\n<p>\nThe final step involves decisions about how to respond to risks. Some risks, which fall below a predetermined risk threshold, can be ignored or accepted. For others, the costs of mitigating the risks may be so high that they outweigh the potential benefits. In such cases, the organization may decide to abandon the corresponding activity. Risks can also be transferred to third parties, such as through insurance that compensates for incurred expenses. The final option is to account for risks when determining strategies and actions that balance costs with potential benefits. This way, the organization will make decisions about implementing strategies to maximize benefits and minimize potential costs.<\/p>\n<p><img decoding=\"async\" alt=\"Structuring Risks and Solutions When Using Big Data for Obtaining Official Statistics\" src=\"\/wp-content\/uploads\/2020\/03\/ff6b25f08f78d1619dcf2c657ae0bb63.jpeg\" style=\"display:block;margin: 0 auto;\" \/><br \/>\n<\/p>\n<h4>2.2. Quality Systems<\/h4>\n<p>\nA target group consisting of representatives from national and international statistical organizations developed a preliminary quality framework for statistics derived from big data in 2014. The target group operated under the UNECE\/HLG project \"The Role of Big Data in Modernizing Statistical Production.\" It expanded existing quality systems designed to assess statistics obtained from administrative data sources, incorporating quality indicators that were deemed relevant for large data sources.<\/p>\n<p>Within this framework, a distinction is made between three phases of the business process: input, processing, and output. The input phase corresponds to the \"design\" and \"collection\" phases of the GSBP, processing relates to the \"process\" and \"analysis\" phases, while the output phase is equivalent to the \"dissemination\" phase.<\/p>\n<p>The structure uses a hierarchical framework derived from the administrative data structure developed by the Statistics Netherlands (7 Daas, P., S. Ossen, R. Vis-Visschers, and J. Arends-Toth, (2009), Checklist for the Quality evaluation of Administrative Data Sources. Statistics Netherlands, The Hague\/Heerlen). The quality measurements are embedded into a hierarchical framework known as hyperspaces. The three defined hyperdimensions are 'source', 'metadata', and 'data'. Quality measurements are embedded within these hyperdimensions and assigned to each stage of production. For the input phase, additional aspects such as 'privacy and confidentiality', 'complexity' (according to the data structure), 'completeness' of metadata, and 'connectivity' (the ability to link data with other data) have been proposed to enhance the standard quality model. For each quality indicator, factors relevant to their description as well as possible indicators are suggested.<\/p>\n<p>In the context of this article, risks can be excluded from these factors. For example, the factors to consider for measuring quality in the 'institutional\/business environment' include the stability of the data provider. The associated risk might be that the data will not be available from the data provider in the future. Another example relates to the recently proposed quality aspect of privacy and security. One of the significant factors is 'perception', meaning the potential negative perception of the intended use of specific data sources by various stakeholders.<\/p>\n<h2>3. Risks Associated with Data Access<\/h2>\n<p><\/p>\n<h4>3.1. Lack of Access to Data<br \/>\n3.1.1. Description<\/h4>\n<p>\nThis risk consists of the project related to the development of BOSP not gaining access to the necessary big data source (BDS).<\/p>\n<p>So far, OSC has learned the hard way that even getting off the starting blocks and obtaining this access can sometimes be an insurmountable obstacle. Sometimes it is easy to gain access to a specific source \u2014 such as call detail records (CDR) \u2014 for testing\/research purposes, but much harder (for legal or commercial reasons) to gain access for production purposes.<\/p>\n<h4>3.1.2. Probability<\/h4>\n<p>\nThe probability largely depends on the characteristics of BDS. If it concerns large administrative data, it may only amount to 1, particularly if (as in the case of traffic loop data studied by Daas et al. 8 Daas, P., M. Puts, B. Buelens, and P. van den Hurk. 2015. \"Big Data as a Source for Official Statistics\". Journal of Official Statistics 31 (2). (Forthcoming; publication foreseen for June 2015.)) there are no personal data protection issues. If the case of BDS pertains to an individual, particularly if it is sensitive (e.g., in terms of data protection) or valuable (from a commercial perspective), the probability may be very high (5).<\/p>\n<h4>3.1.3. Impact<\/h4>\n<p>\nThe impact depends on BOSP and how BDS is utilized. If BDS is central, the impact may be very high (4 = it is generally impossible to produce BOSP), while it may be lower if it is still possible to produce BOSP (albeit with lower quality) relying on other URBs, leading to an impact in the range of 2-3.<\/p>\n<h4>3.1.4. Prevention<\/h4>\n<p>\nTo reduce the risk of loss of access, preliminary contacts should be established with the data provider and a long-term access agreement to the data should be concluded. Additionally, a comprehensive legal analysis should be conducted concerning the specific combination of BDS and BOSP. The possibilities for accessing data should also be evaluated under current or future legislation.<\/p>\n<h4>3.1.5. Mitigation<\/h4>\n<p>\nIf there are alternative BDS that can be used for BOSP, they could be explored instead. If there is no way to produce BOSP without BDS, and if it is impossible to overcome the lack of access, efforts should cease, and a new BOSP will not come to fruition.<\/p>\n<h4>3.2. Loss of Data Access<br \/>\n3.2.1. Description<\/h4>\n<p>\nThis risk involves statistical management losing the BDS underlying BOSP.<\/p>\n<h4>3.2.2. Probability<\/h4>\n<p>\nIf the BOSP is already in production, there is usually a certain stability, and in some cases, the risk may be very low (1). However, in particular, in the case of private entities with insufficiently solid agreements, nothing prevents, for instance, new management from changing data provision policies, leading to a moderate risk of disruption (3). Moreover, if the BDS is linked to unstable activities, there is always a risk that the provider may simply go bankrupt, and the risk may even be higher (4).<\/p>\n<h4>3.2.3. Impact<\/h4>\n<p>\nSince the existing BOSP may be impossible to produce, there is often a very strong impact (5). In other cases, when the BDS is ancillary, the impact may be more about a loss of quality with an effect ranging between 2-3.<\/p>\n<h4>3.2.4. Prevention<\/h4>\n<p>\nThe prevention strategy is similar to the strategy of lack of access to data but with an increased focus on continuous vigilance even in production conditions.<\/p>\n<p>Not putting all your eggs in one basket (i.e., having multiple BDS underpinning each BSOP) can also be a strategy, but this may be either impractical or too costly.<\/p>\n<h4>3.2.5. Mitigation<\/h4>\n<p>\nIf the BDS is the result of unstable activity, a new BDS reflecting the same social phenomenon may gradually become available. However, it would be too late to start \"market scanning\" as soon as the BSOP fails; continuous vigilance will be required \u2014 which can be difficult to achieve.<\/p>\n<h2>4. Risk Associated with the Legal Environment<\/h2>\n<p><\/p>\n<h4>4.1. Non-compliance with Relevant Legislation<br \/>\n4.1.1. Description<\/h4>\n<p>\nThis risk consists of a project related to the development of BOSP, which does not take into account the relevant legislation, making the BOSP non-compliant with the specified legislation. This may pertain to data protection laws, regulations concerning the burden of response, etc.<\/p>\n<h4>4.1.2. Probability<\/h4>\n<p>\nGiven the OSC's lack of awareness about big data, it is not improbable that accidental (3) non-compliance may occur. The probability is generally related to the BDS, as the less \"sensitive\" the source, the lower the likelihood of non-compliance.<\/p>\n<h4>4.1.3. Impact<\/h4>\n<p>\nThe impact is generally critical (4) in the sense that non-compliant production will require halting the BOSP (or, if it has not yet reached the implementation stage, its development must be terminated). This can even be extreme (5), as the reputational risks arising from non-compliant (\u2018illegal\u2019) official statistics may have consequences.<\/p>\n<h4>4.1.4. Prevention<\/h4>\n<p>\nA thorough legal analysis is required for any BOSP \u2014 and this occurs in several stages (what is acceptable at the development\/exploration stage may not be the same at the implementation\/production stage). This, in turn, may lead to re-engineering the BOSP to ensure compliance.<\/p>\n<h4>4.1.5. Mitigation<\/h4>\n<p>\nDepending on the severity of the non-compliance, the first step may be to transition the BOSP to offline mode.<\/p>\n<p>Re-engineering the BOSP for compliance may be an option, but whether the BOSP will be \u2018saved\u2019 in this way greatly depends on the nature of the non-compliance.<\/p>\n<h4>4.2. Adverse Changes in the Legal Environment<br \/>\n4.2.1. Description<\/h4>\n<p>\nNew legislation may be introduced relating to the BOSP being developed, effectively rendering the BOSP non-compliant.<\/p>\n<h4>4.2.2. Probability<\/h4>\n<p>\nIt is possible that advocates for enhanced data protection will succeed in introducing new requirements that will directly or indirectly affect the ability to create specific BOSP. A probability in the range of 2-3 seems like a realistic assessment.<\/p>\n<h4>4.2.3. Impact<\/h4>\n<p>\nThe impact is generally critical (4), in the sense that non-compliant production will require stopping the BOSP.<\/p>\n<h4>4.2.4. Prevention<\/h4>\n<p>\nCertain business information should be regularly monitored to track legislative developments \u2014 possibly also to influence it by advocating for official statistics at relevant (e.g., advisory) forums.<\/p>\n<h4>4.2.5. Mitigation<\/h4>\n<p>\nProvided that proactive monitoring has been conducted, there may be time for re-engineering the BOSP to align it with the new legislation from the day it comes into effect.<\/p>\n<p>If, on the other hand, monitoring has not been conducted, so that the new legislation has become a 'surprise' \u2014 or if the legislation is so radical that there is no way to make the BOSP incompatible \u2014 the only option may be to disable the BOSP.<\/p>\n<h2>5. Risks Related to Privacy and Data Security<\/h2>\n<p><\/p>\n<h4>5.1. Data Security Breaches<br \/>\n5.1.1. Description<\/h4>\n<p>\nThis risk relates to unauthorized access to data stored in statistical management systems. Third parties may gain access to embargoed data, for example, due to the release of a schedule (9 For any BOSP that is completely based on a single BDS, it is inevitable that the data will implicitly be known to the original data owner, and if the methodology is transparent, the derived statistics will also be known. This situation is not addressed here, but rather in the risk related to the misuse of power by owners.) (10 Additionally, this data may carry a risk of privacy violations. This risk will be considered separately.). This may include, for example, data that investors expect on the stock market.<\/p>\n<h4>5.1.2. Likelihood<\/h4>\n<p>\nRegarding the technical aspects of protecting the IT environment in the statistical department, the risk is the same for BDSs as it is for traditional sources. However, there are two additional aspects that need to be considered.<\/p>\n<p>Firstly, with some BDS, the overall risk slightly increases due to the fact that the data security of the original owner may be compromised. This could be related, for example, to industrial espionage or hacking.<\/p>\n<p>Secondly, once potentially valuable data begins to be stored in the office, the risk of attracting malicious intentions will increase. If the stored data has a very high value for the business, one should be prepared for a very high likelihood of attacks directed at the IT infrastructure, therefore the likelihood of a breach may be potentially higher (4).<\/p>\n<p>If the stored data is not perceived as valuable, the overall likelihood appears to be not very high \u2014 from (1) to (3) depending on the data source.<\/p>\n<h4>5.1.3. Impact<\/h4>\n<p>\nThe potential reputational damage can be significant (5). What is crucial in the case of BDS is that if a security breach occurs at the original owner, the impact on the reputation of the statistical authority is expected to be lower than if the breach happened with the data held by it.<\/p>\n<p>On the other hand, it is possible that a breach in the statistical authority can have negative consequences for the original owner. In this case, there could again be a severe negative impact due to the damage in terms of trust between the provider and the statistical authority (5).<\/p>\n<h4>5.1.4. Prevention<\/h4>\n<p>\nWhat characterizes the BDS case is that the security procedures of the original owner may be relevant. It is unlikely that statistical authorities will receive auditing powers to oversee this. Owners whose data is used to create records with sensitive publication schedules should be informed about the implications for official statistics of a potential security breach in their premises and should receive an official guarantee that appropriate security procedures are in place.<\/p>\n<p>A direct way to prevent a significant impact of a security breach at the owner's facility on the statistical authority is to ensure that multiple sources are used for the same product, so that one compromised source is insufficient to determine the final figure. The advantage of this approach is that greater control remains in the hands of the statistical authority.<\/p>\n<p>The way to prevent the negative consequences of security breaches in the statistical office for the original data owner is to find a working method that does not involve transmitting potentially sensitive data from the owner's perspective to statistical management. In its raw form. A possible preventive approach is the use of aggregated data. However, it should be noted that some forms of aggregation, such as those aimed at preventing the identification of individual population members, may not be suitable in this case. One reason for this could be that the risk to the owner is related to the commercial value of the data, which may be significant even after achieving anonymity.<\/p>\n<h4>5.1.5. Mitigation<\/h4>\n<p>\nIn the event of a data breach under the management of statistical authorities, the mitigation measures will be the same as for traditional sources, unless there has been a negative impact on the original owner.<\/p>\n<p>In the case of negative consequences for the original owner, the statistical authority must review and strengthen its security procedures and clearly communicate and demonstrate its commitment to this.<\/p>\n<p>If the breach occurred within the premises of the original owner, the relevant statistical service must clearly communicate the situation and insist on improving the owner's security procedures. An alternative provider may be sought if necessary.<\/p>\n<h4>5.2. Data Privacy Breaches<\/h4>\n<p><\/p>\n<h4>5.2.1. Description<\/h4>\n<p>\nThis is the risk that the privacy of one or more individuals from the statistical population will be violated. This may be related to attacks on IT infrastructure due to pressure from other government agencies or due to inadequate controls over the disclosure of statistical data.<\/p>\n<h4>5.2.2. Likelihood<\/h4>\n<p>\nAs with the risk of data security breaches, the technical conditions for storing microdata do not change significantly with the addition of BDS. However, there are caveats here as well.<\/p>\n<p>Microdata from certain data sources may have high business value, so their storage increases the likelihood of attacks.<\/p>\n<p>In addition, some microdata can be potentially very useful for other government agencies, such as law enforcement, taxation, or healthcare. Under certain circumstances, the commitment to the principle of statistical confidentiality may come under considerable pressure.<\/p>\n<p>Regarding failures in the control of statistical information disclosure, an established practice already exists. BDS may allow for statistics to be produced for small subgroups of the population or provide the possibility of linking aggregated data from different BDS, which may increase the likelihood of risks. Moreover, new sources will, however, require new methodological developments, so the real danger is that the methodology for controlling information disclosure is not updated properly.<\/p>\n<p>Overall, with reasonable preventive measures, the probability can be maintained at reasonable levels, but since there are many different and diverse factors, the appropriate assessment here seems to be that the probability is high (4).<\/p>\n<h4>5.2.3. Impact<\/h4>\n<p>\nThe potential reputational damage can be significant (5). Similar to the risk of data security breaches, problems in statistical governance can have negative consequences for the original owner. The impact of such an event could potentially be even greater, especially if current trends in public opinion persist. Damage in the relationship between the data provider and statistical governance is also expected to be quite substantial.<\/p>\n<h4>5.2.4. Prevention<\/h4>\n<p>\nThe infallible way to prevent such risks is to avoid having microdata from BDS altogether (although storing other microdata still carries its associated risks, albeit with different probabilities and impacts). This approach, much like the risk of data breaches, necessitates the development of alternative methods for utilizing data for statistical purposes. Moreover, the varying nature of sources will mean that new methodologies must be developed to balance the competing objectives of extracting as much useful information as possible while ensuring privacy protection from potential threats.<\/p>\n<p>When storing microdata, IT security mechanisms and access controls must be at the required level and monitored continuously. Special attention must be given to the security of new data acquisition methods. Ironically, a new method could be the physical transportation of storage devices (e.g., hard drives). If this method is used, delivery must be physically secured, and encryption must be employed.<\/p>\n<h4>5.2.5. Mitigation<\/h4>\n<p>\nMitigation measures here are essentially the same as in the case of a data breach. If the breach is caused by pressure from another government entity, one should seize the opportunity to strengthen governance independence so that similar breaches become even more difficult in the future.<\/p>\n<h4>5.3. Data Source Manipulations<br \/>\n5.3.1. Description<\/h4>\n<p>\nThird-party data providers, such as social media data or voluntarily provided data, are at risk of manipulation. This can be done either by the data provider themselves or by third parties. For instance, many false messages on social media may be generated to push a statistical index computed from this data in one direction or another if it's known that the index is derived from such data.<\/p>\n<p>For voluntarily provided data, there may be cases where volunteers represent a specific interest group with a particular agenda.<\/p>\n<h4>5.3.2. Probability<\/h4>\n<p>\nFor data that can be manipulated for significant benefit, the likelihood is higher. This could be data for which statistics are of interest, such as the stock market. In light of recent scandals associated with LIBOR and Forex, it can be assumed that as long as there is an incentive, attempts to manipulate data will be probable.<\/p>\n<p>For statistics based on voluntarily provided data, one need only look at the recent PR practice of hiring individuals to pretend to have certain opinions and being paid for their public expression (e.g., on internet forums) to conclude that the likelihood is not insignificant. Overall, a figure between 3 and 4 seems adequate.<\/p>\n<h4>5.3.3. Impact<\/h4>\n<p>\nA major problem with manipulations is that they can last a long time without detection. If manipulations persist for an extended period, their impact on quality can become significant. Additionally, the damage to public trust in official statistics can also be considerable, especially if the role of statistical agencies as providers of quality data is publicly emphasized. On the other hand, if manipulations are detected in time and subsequently published, it can actually improve public perception. Except in extremely bad cases, a maximum impact can be imagined (3).<\/p>\n<h4>5.3.4. Prevention<\/h4>\n<p>\nConducting regular control exercises with alternative sources is one possible preventive approach. These alternative sources can be traditional or otherwise. Using statistics based on a combination of sources can help prevent significant impacts of manipulations. In cases where provider-initiated manipulations are feared, legal agreements can also be one way to prevent such practices.<\/p>\n<h4>5.3.5. Mitigation<\/h4>\n<p>\nFrom a public relations damage perspective, the mitigating measures to be taken here are little different from those for addressing any crisis.<\/p>\n<p>From the perspective of data quality, it would be helpful if past data could be corrected in such a way that even with significant delays, the correct series could be provided.<br \/>\nThis is produced. Regular benchmarking can be useful for this purpose. Note that the goal of comparative analysis in this case is somewhat different from the goal of prevention. For prevention, it is important to quickly notice and investigate any suspicious discrepancies between benchmark test data and BDS. Old useful data is always beneficial for mitigating consequences.<\/p>\n<p>Moreover, care should be taken to prevent similar manipulations in the future \u2014 in especially delicate cases, this may mean obtaining potentially excessive data from multiple suppliers for comparative analysis.<\/p>\n<h4>5.4. Unfavorable Public Perception of the Use of Big Data by Official Statistics<br \/>\n5.4.1. Description<\/h4>\n<p>\nThe media and the general public are very sensitive to issues of privacy and the use of personal data from large data sources, especially in the context of the secondary use of data by government bodies taking administrative or legal action against citizens. An unfavorable use may be the positioning of speed control based on navigation data analysis (11 See <noindex><a rel=\"nofollow\" href=\"http:\/\/www.theguardian.com\/technology\/2011\/apr\/28\/tomtom-satnav-data-police-speed-traps\">www.theguardian.com\/technology\/2011\/apr\/28\/tomtom-satnav-data-police-speed-traps<\/a><\/noindex>). <br \/>\nThe specific case of TomTom Netherlands led to a significant drop in demand for TomTom devices and prompted the company to limit access to data. In this specific case, the data pertained to individuals but indicated speed levels on stretches of road.<\/p>\n<p>However, there may be big data applications that are positively perceived by the public. One example can be applications that prevent crimes such as burglary based on big data methods.<\/p>\n<p>Both positive and negative public opinion can have a strong influence on the use of BDS in the context of producing official statistics.<\/p>\n<p>The consequence of negative public perception may be that:<\/p>\n<ul>\n<li>BDS will no longer be available for statistical offices, either due to decisions by data providers or government decisions not to use the data, or<\/li>\n<li>the use of data will be limited, which may hinder production if certain BOSP.<\/li>\n<\/ul>\n<p><\/p>\n<h4>5.4.2. Probability<\/h4>\n<p>\nFactors that may influence the likelihood of such an event or its impact on statistical production:<\/p>\n<ul>\n<li>data confidentiality, i.e., how easily individuals can be identified;<\/li>\n<li>the volume of information that the data reveals about individuals, for instance, increased by linking data from different sources;<\/li>\n<li>the type of data, for example, financial transactions are perceived as more confidential than other data;<\/li>\n<li>the type of potential action that may be taken against citizens, such as fining individuals for speeding;<\/li>\n<li>a vague legal environment in which data providers and users operate, or when legal conditions conflict with social ethical opinions\/standards;<\/li>\n<li>the degree of dependence on a specific data source for obtaining statistics; at the exploratory stage, this factor may have little significance. However, it can greatly affect the acquisition of statistics at a later stage and therefore should also be considered at the exploratory stage. One of the issues may be that the final extent of data usage is initially unknown, as data sources can potentially serve more than one statistical area.<\/li>\n<\/ul>\n<p>\nEstimating the timing of undesirable phenomena is impossible, as public mobilization is often initiated by media coverage of events that negatively impact citizens. Nevertheless, with the increasing use of big data by governments and private enterprises, and particularly with active marketing of data for purposes other than those that led to their initial collection, such events are more likely to occur.<\/p>\n<p>Events that significantly influence public perception are not frequent but rather random (3) and distant (2). As the use of large data sources increases, the likelihood also rises.<\/p>\n<h4>5.4.3. Impact<\/h4>\n<p>\nThe impact of the event greatly depends on the factors discussed above. Overall, the influence is more serious for already established statistical data production, as it may be necessary to halt certain actions. The impact also depends on the availability of alternative data sources, although it may happen that public perception does not differentiate between different data sources in the event of materialization. In the current state of big data usage, it seems that these sources cannot completely replace traditional data sources but rather complement the existing statistics. This will diminish the impact of events. Therefore, the impact of the event is considered to range from 2 (minor) to 3 (major). At the production stage, the impact can increase to 4 (critical value).<\/p>\n<h4>5.4.4. Prevention<\/h4>\n<p>\nPreventive measures may include establishing ethical principles for big data in official statistics. Ethical guidelines should be based on principles such as the code of practice for European statistics or the fundamental principles of official statistics (12 <noindex><a rel=\"nofollow\" href=\"http:\/\/unstats.un.org\/unsd\/dnss\/gp\/fundprinciples.aspx\">unstats.un.org\/unsd\/dnss\/gp\/fundprinciples.aspx<\/a><\/noindex>). The next measure will be to establish a communication strategy that will publish the results of ethical guidelines for the public and can be used to inform stakeholders about the ethical use of BDS for BOSP.<\/p>\n<p>A separate risk assessment for a specific BDS can be conducted to identify risks and propose preventive or mitigating actions based on ethical principles. A separate risk assessment may also include stakeholders such as data protection agencies to ensure the identification of all risks and the justification of actions.<\/p>\n<h4>5.4.5. Mitigation<\/h4>\n<p>\nThe communication strategy must also include measures in case of growing negative public sentiment. A separate risk assessment should collect positive examples of data usage and measures to prevent data misuse that may need to be mandated at the political level, and the statistical community may find it challenging to influence them effectively.<\/p>\n<h4>5.5. Loss of Trust \u2014 not resulting from observation<br \/>\n5.5.1. Description<\/h4>\n<p>\nUsers of official statistics generally have a high level of confidence in the accuracy and reliability of statistical data. This is based on the fact that the production of statistical data is embedded in a reliable and publicly available methodological framework, along with documentation regarding the quality of the statistical product. Additionally, most statistical data is based on observations, i.e., obtained from surveys or censuses that establish an easily understandable connection between observation and statistical data. The use of BDS, which are not collected for the primary purpose of statistics, carries the risk that these relationships will be lost, leading users to lose trust in official statistical data. An example related to the last round (2010) of the census involves cases in some countries where statistical data were obtained using multiple sources and statistical models. In several instances, stakeholders contested the statistical data.<\/p>\n<h4>5.5.2. Probability<\/h4>\n<p>\nThe probability of risk occurrence depends on factors such as the complexity of the statistical\/methodological model, the reliability of relationships between BSD and BOSP, or the alignment with other statistical data. The probability should range from 3 (random) to 4 (likely), indicating that it may occur several times or often.<\/p>\n<h4>5.5.3. Impact<\/h4>\n<p>\nThe impact of the risk's occurrence will largely depend on whether NSOs can successfully demonstrate the accuracy and reliability of the statistical data. If this cannot be achieved, the repercussions in terms of loss of trust and credibility may also affect other statistical areas, calling into question not only certain statistical data but also the organization itself. NSOs would lose their competitive advantage over other private entities operating in this field.<\/p>\n<h4>5.5.4. Prevention<\/h4>\n<p>\nPreventive actions will involve the development and publication of a scientifically grounded methodology recognized by the scientific community, enriching the data with quality metadata, ensuring the consistency of BOSP with non-BOSP, and implementing rigorous quality control.<\/p>\n<p>Before commencing statistical production, BOSP could be published as experimental, and stakeholders would be encouraged to challenge BOSP to validate or enhance it.<\/p>\n<h4>5.5.5. Mitigation<\/h4>\n<p>\nThere are two cases to distinguish. In the event that the statistical data is disputed but possesses high\/sufficient quality (correct\/accurate), it would be enough to explain and convey the statistical data to the public, providing easily understandable examples.<\/p>\n<h2>6. Risks Related to Skills<\/h2>\n<p><\/p>\n<h4>6.1. Lack of Specialists<br \/>\n6.1.1. Description<\/h4>\n<p>\nAnalyzing the digital traces left by individuals during their activities requires specific data analysis tools that are not currently common in official statistics. Firstly, using indirect data about people's activities instead of direct surveys demands the application of statistical models and, consequently, skills in reasoning and machine learning. Secondly, these digital records consist of data that often lack the conventional table format typical of survey results, with rows corresponding to statistical units and columns detailing specific characteristics of those units. Digital traces are also presented in forms of text, sound, images, and video. Extracting relevant statistical information from these data types requires expertise in natural language processing, audio signal processing, and image processing. Thirdly, these data sources tend to provide large datasets, the processing of which necessitates a solid understanding of distributed computing methodologies.<\/p>\n<p>The risk of lacking experts lies in obtaining data from one of these new large data sources, as statistical offices do not have the capacity to process and analyze it properly due to their staff lacking the necessary skills.<\/p>\n<h4>6.1.2. Probability<\/h4>\n<p>\nThe probability of this risk will depend on three factors: 1) the specific types of skills required for each type of big data source and the likelihood that the statistical office will find an opportunity to explore such a source; 2) the current availability of necessary skills within the statistical office; and 3) the organizational culture of the statistical office.<\/p>\n<p>When it comes to the types of skills that may be required, it should be noted that not all sources demand all the skills listed above. Some (like Google Trends data) do not require distributed computing as they are already pre-processed by the data holder or involve signal processing skills, primarily requiring statistical modeling skills. However, there is a great variety of big data sources for which distributed computing, signal processing, and machine learning skills are generally needed. At the same time, proper research into these digital footprints will require processing multiple sources. Thus, there is a high likelihood that large data sources becoming available for statistical management will necessitate these unusual skills, and the probability of this risk is very high (5).<\/p>\n<p>As for the current availability of the necessary skills, this will depend on the specific statistical management. Even if the survey methodology is less widespread than the survey methodology, it is still used in official statistics in certain fields. Therefore, even if some redistribution of human resources may be required, statistical managements can find solutions on their own. Regarding distributed computing skills, mainly associated with IT, they will depend on how the IT infrastructure is managed within the organization. Depending on how aligned the IT department is, solutions may be found within the context of existing agreements. Nonetheless, signal processing and machine learning skills generally do not exist in most official statistical managements, and the application of these skills cannot be outsourced, as they must be applied by experts in statistics. Consequently, from this perspective, the probability of this risk also seems very high (5).<\/p>\n<p>Organizational culture will also influence the likelihood of this risk. Having personnel willing to acquire necessary skills through self-learning can give the organization the ability to respond to situations with new data sources requiring skills different from the usual. This will depend on the organizational culture of statistical management, particularly whether it encourages employees to learn new skills and allows them time for self-study.<\/p>\n<p>Thus, the probability that statistical management will be unable to process and analyze new data sources due to a lack of skills among its employees will be between likely (4) and frequent (5), depending on the organization's self-learning culture.<\/p>\n<h4>6.1.3. Impact<\/h4>\n<p>\nStatistical management unable to process and analyze large data sources due to a lack of skills among its employees may face two potential negative consequences: 1) the data source will not be studied, at least not in full; 2) the source will be misused.<\/p>\n<p>The lack of capability to fully explore the potential of a valuable large data source will have a minor impact (2) in the short term, as statistical management indeed has statistical tools to meet current needs. However, in the long term (and perhaps even in the medium term), the consequences of losing this opportunity will be critical (4), as statistical management increasingly faces competition from private providers who do not have the same institutional structure to guarantee the independence of statistical data.<\/p>\n<p>However, improper use of the source will have extremely negative consequences for statistical agencies, as official statistics largely depend on their reputation in fulfilling their mission. Nevertheless, we can argue that the most crucial skill, if overlooked, can lead to incorrect results is statistical inference, particularly model-based inference, which is also less likely to be absent. Therefore, the expected impact will be more critical (4) than extreme.<\/p>\n<h4>6.1.4. Prevention<\/h4>\n<p>\nStatistical agencies can actively prevent this risk in two ways: 1) training; and 2) recruitment.<\/p>\n<p>Statistical agencies can provide staff with the necessary skills by clearly defining the skills needed to utilize large data sources in statistical production, compiling a list of existing staff skills, identifying training needs, and then organizing training courses.<\/p>\n<p>Statistical agencies can also recruit new staff with the required skills. This seems to have serious limitations, as statistical agencies may not be able to recruit a critical mass of personnel for a situation where the use of large data sources becomes widespread in the department, and new employees will still require several years to reach the expertise level of existing staff. However, at least some of the new employees hired through regular personnel updates may possess skills related to big data.<\/p>\n<h4>6.1.5. Mitigation<\/h4>\n<p>\nWhen faced with a situation where new sources of big data are available without staff with the necessary skills, statistical agencies can mitigate the negative consequences in two ways: 1) subcontracting; and 2) collaboration.<\/p>\n<p>Statistical agencies may enter into contracts for data processing and the analysis of new sources of big data with other organizations that provide these services. This appears to be a viable solution, as a new sector of enterprises specializing in processing such data is emerging. However, this solution comes with certain risks, as statistical agencies will have less control over the production of potentially sensitive statistical products. Additionally, this approach has the drawback of not allowing employees of statistical agencies to learn and acquire necessary skills.<\/p>\n<p>Collaboration with other organizations that have employees with the necessary skills and are also interested in exploring big data sources appears to be a more promising solution. This collaboration can take the form of joint projects with employees of statistical agencies and staff from other organizations on equal terms, sharing their knowledge. This would not only help mitigate the risk of skill shortages but also enable staff at statistical agencies to gain these skills.<\/p>\n<h4>6.2. Loss of Experts to Other Organizations<br \/>\n6.2.1. Description<\/h4>\n<p>\nThis risk involves statistical agencies losing their personnel to other organizations after they have acquired skills related to big data.<\/p>\n<h4>6.2.2. Probability<\/h4>\n<p>\nThe probability of this risk will depend on two factors: 1) the existence of attractive opportunities in organizations outside of official statistics; 2) the working conditions in statistical agencies.<\/p>\n<p>When it comes to opportunities in organizations outside official statistics, the likelihood of this risk appears plausible (4). There is a high demand for individuals with skills in big data in the private sector, as well as in other public sector organizations. After acquiring big data skills, official statistics will gain a comparative advantage by being experienced professionals in the field of statistics. In addition to specific big data skills, other organizations require data specialists who possess more traditional skills, such as assessing user needs and developing key performance indicators (KPIs), which are common for official statisticians. Furthermore, it is expected that employees who are more prone to acquiring new skills will also be those who are more open to career changes and will leave statistical offices.<\/p>\n<p>Regarding working conditions in statistical offices, this will obviously depend mainly on the specific office. Nevertheless, statistical offices still offer attractive professional opportunities for quantitatively inclined individuals. Statistical offices provide the widest range of possible domains to work in and the largest selection of data to work with. This somewhat reduces the likelihood that statistical offices will lose their personnel due to unforeseen circumstances (3).<\/p>\n<h4>6.2.3. Impact<\/h4>\n<p>\nThe impact of this risk will be the same as the risk of not having personnel with the relevant skills in the first place. Therefore, the impact will be critical (4), as mentioned above.<\/p>\n<h4>6.2.4. Prevention<\/h4>\n<p>\nIt seems that the only way for statistical agencies to mitigate this risk is to provide attractive working conditions for their employees. This is generally true for all staff. However, in the specific case where employees are open to acquiring new skills, particularly in handling big data, working conditions can be improved by offering them training opportunities where they can develop their professional interests. Statistical agencies can also focus on being open to new innovative projects and ideas related to new sources of big data, coming from statisticians working in various fields of statistics. Finally, preventing the loss of personnel to other organizations based on their big data skills will depend on good identification of staff capable and willing to work with such data, and on providing good opportunities for their professional development.<\/p>\n<h4>6.2.5. Mitigation<\/h4>\n<p>\nThis risk will be mitigated in relation to the lack of personnel with the necessary skills: 1) subcontracting; and 2) collaboration.<\/p>\n<h2>7. Discussion<\/h2>\n<p>\nFrom this initial review, it is clear that it is impossible to establish a single probability or impact for the 'big data risk' \u2014 in general, both metrics depend significantly on the source of big data, as well as on 'official statistics based on big data.'<br \/>\nproduct.<\/p>\n<p>Thus, we conclude that the logical next step in this direction is to undertake a series of possible pilot projects (each involving a combination of one or more BDSs and one or more BDOSs) as a starting point, and \u2014 for each such pilot \u2014 to strive to assess the likelihood and impact of each risk.<\/p>\n<p>To this end, we are on the verge of launching a stakeholder survey aimed at assessing the OSC's evaluation of the probability, impact (and possible actions for prevention \/ mitigation) regarding a number of possible pilot projects \u2014 and to request OSC suggestions regarding risks that we have not included in this document.<\/p>\n<p><b class=\"spoiler_title\">8. REFERENCES<\/b>UNECE (2014), \"A suggested Framework for the Quality of Big Data\", Deliverables of the UNECE Big Data Quality Task Team, <noindex><a rel=\"nofollow\" href=\"http:\/\/www1.unece.org\/stat\/platform\/download\/attachments\/108102944\/Big%20Dat\">www1.unece.org\/stat\/platform\/download\/attachments\/108102944\/BigDat<\/a><\/noindex><br \/>\na%20Quality%20Framework%20-%20final-%20Jan08-2015.pdf?version=1&amp;modificationDate=1420725063663&amp;api=v2 <\/p>\n<p>UNECE (2014), \"How big is Big Data? Exploring the role of Big Data in Official Statistics\", <noindex><a rel=\"nofollow\" href=\"http:\/\/www1.unece.org\/stat\/platform\/download\/attachments\/99484307\/Virtual%20Sprint%20Big%20Data%20paper.docx?version=1&amp;modificationDate=1395217470975&amp;api=v2\">www1.unece.org\/stat\/platform\/download\/attachments\/99484307\/VirtualSprintBigDatapaper.docx?version=1&amp;modificationDate=1395217470975&amp;api=v2<\/a><\/noindex><\/p>\n<p>Daas, P., S. Ossen, R. Vis-Visschers, and J. Arends-Toth, (2009), Checklist for the Quality evaluation of Administrative Data Sources, Statistics Netherlands, The Hague\/Heerlen <\/p>\n<p>Dorfman, Mark S. (2007), Introduction to Risk Management (e ed.), Cambridge, UK, Woodhead-Faulkner, p. 18, ISBN 0-85941-332-22)<\/p>\n<p>Eurostat (2014), \"Accreditation procedure for statistical data from non-official sources\" in Analysis of Methodologies for using the Internet for the collection of information society and other statistics, <noindex><a rel=\"nofollow\" href=\"http:\/\/www.cros-portal.eu\/content\/analysismethodologies-using-internet-collection-information-society-and-other-statistics-1\">www.cros-portal.eu\/content\/analysismethodologies-using-internet-collection-information-society-and-other-statistics-1<\/a><\/noindex><\/p>\n<p>Reimsbach-Kounatze, C. (2015), \u201cThe Proliferation of \u201cBig Data\u201d and Implications for Official Statistics and Statistical Agencies: A Preliminary Analysis\u201d, OECD Digital Economy Papers, No. 245, OECD Publishing. <noindex><a rel=\"nofollow\" href=\"http:\/\/dx.doi.org\/10.1787\/5js7t9wqzvg8-en\">dx.doi.org\/10.1787\/5js7t9wqzvg8-en<\/a><\/noindex><\/p>\n<p>Reis, F., Ferreira, P., Perduca, V. (2014) \"The use of web activity evidence to increase the timeliness of official statistics indicators\", paper presented at IAOS 2014 conference, <noindex><a rel=\"nofollow\" href=\"https:\/\/iaos2014.gso.gov.vn\/document\/reis1.p1.v1.docx\">iaos2014.gso.gov.vn\/document\/reis1.p1.v1.docx<\/a><\/noindex><\/p>\n<p>Even if not explicitly mentioning risks, this paper in fact approaches the many risks associated with the use of web activity data for official statistics. Eurostat (2007), Handbook on Data Quality Assessment Methods and Tools, <noindex><a rel=\"nofollow\" href=\"http:\/\/ec.europa.eu\/eurostat\/documents\/64157\/4373903\/05-Handbook-ondata-quality-assessment-methods-and-tools.pdf\/c8bbb146-4d59-4a69-b7c4-218c43952214\">ec.europa.eu\/eurostat\/documents\/64157\/4373903\/05-Handbook-ondata-quality-assessment-methods-and-tools.pdf\/c8bbb146-4d59-4a69-b7c4-218c43952214<\/a><\/noindex><\/p>\n<p>Source: <a content=\"nofollow\" rel=\"nofollow\" href=\"https:\/\/habr.com\/ru\/post\/494044\/\">habr.com<\/a> <\/p>","protected":false,"gt_translate_keys":[{"key":"rendered","format":"html"}]},"excerpt":{"rendered":"<p>\u041f\u0440\u0435\u0434\u0438\u0441\u043b\u043e\u0432\u0438\u0435 \u043f\u0435\u0440\u0435\u0432\u043e\u0434\u0447\u0438\u043a\u0430 \u041c\u0430\u0442\u0435\u0440\u0438\u0430\u043b \u0437\u0430\u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043e\u0432\u0430\u043b \u043c\u0435\u043d\u044f, \u0432 \u043f\u0435\u0440\u0432\u0443\u044e \u043e\u0447\u0435\u0440\u0435\u0434\u044c \u0438\u0437-\u0437\u0430 \u0442\u0430\u0431\u043b\u0438\u0446\u044b \u043d\u0438\u0436\u0435: \u0421 \u0443\u0447\u0435\u0442\u043e\u043c \u0442\u043e\u0433\u043e, \u0447\u0442\u043e \u0441\u0442\u0430\u0442\u0438\u0441\u0442\u0438\u043a\u0438 (\u0430 \u0440\u043e\u0441\u0441\u0438\u0439\u0441\u043a\u0438\u0435, \u043d\u0430 \u0433\u0435\u043d\u0435\u0442\u0438\u0447\u0435\u0441\u043a\u043e\u043c \u0443\u0440\u043e\u0432\u043d\u0435), \u043c\u044f\u0433\u043a\u043e \u0433\u043e\u0432\u043e\u0440\u044f, \u043d\u0435 \u043b\u044e\u0431\u044f\u0442 \u0432\u0441\u0435, \u0447\u0442\u043e \u043e\u0442\u043b\u0438\u0447\u0430\u0435\u0442\u0441\u044f \u043e\u0442 \u043b\u0438\u043d\u0435\u0439\u043d\u043e\u0439 \u0437\u0430\u0432\u0438\u0441\u0438\u043c\u043e\u0441\u0442\u0438, \u044d\u0442\u0438 \u043f\u0430\u0440\u043d\u0438 \u0443\u043c\u0443\u0434\u0440\u0438\u043b\u0438\u0441\u044c \u043f\u0440\u043e\u0442\u0430\u0449\u0438\u0442\u044c \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u0435 \u0444\u0443\u043d\u043a\u0446\u0438\u0438 \u0430\u043a\u0442\u0438\u0432\u0430\u0446\u0438\u0438 \u0432 \u043f\u0430\u0440\u0430\u0431\u043e\u043b\u0438\u0447\u0435\u0441\u043a\u043e\u043c \u0432\u0438\u0434\u0435 \u0434\u043b\u044f \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0435\u043d\u0438\u044f \u0441\u0442\u0435\u043f\u0435\u043d\u0438 \u0440\u0438\u0441\u043a\u0430 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u044f BigData \u0432 \u043e\u0444\u0438\u0446\u0438\u0430\u043b\u044c\u043d\u043e\u0439 \u0441\u0442\u0430\u0442\u0438\u0441\u0442\u0438\u043a\u0435. \u041c\u043e\u043b\u043e\u0434\u0446\u044b. \u0415\u0441\u0442\u0435\u0441\u0442\u0432\u0435\u043d\u043d\u043e, \u0441\u0442\u0430\u0442\u0438\u0441\u0442\u0438\u043a\u0438 \u0434\u043e\u0431\u0430\u0432\u0438\u043b\u0438 \u0441\u0432\u043e\u0435 [&hellip;]<\/p>\n","protected":false,"gt_translate_keys":[{"key":"rendered","format":"html"}]},"author":1,"featured_media":75637,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[688],"tags":[],"class_list":["post-75636","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-administrirovanie"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.2 - aioseo.com -->\n\t<meta name=\"description\" content=\"\u041f\u0440\u0435\u0434\u0438\u0441\u043b\u043e\u0432\u0438\u0435 \u043f\u0435\u0440\u0435\u0432\u043e\u0434\u0447\u0438\u043a\u0430 \u041c\u0430\u0442\u0435\u0440\u0438\u0430\u043b \u0437\u0430\u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043e\u0432\u0430\u043b \u043c\u0435\u043d\u044f, \u0432 \u043f\u0435\u0440\u0432\u0443\u044e \u043e\u0447\u0435\u0440\u0435\u0434\u044c \u0438\u0437-\u0437\u0430 \u0442\u0430\u0431\u043b\u0438\u0446\u044b \u043d\u0438\u0436\u0435:\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Yuri Gagarin\"\/>\n\t<link rel=\"canonical\" href=\"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.2\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"ProHoster | \u041a\u0443\u043f\u0438\u0442\u044c \u043d\u0430\u0434\u0435\u0436\u043d\u044b\u0439 \u0445\u043e\u0441\u0442\u0438\u043d\u0433 \u0434\u043b\u044f \u0441\u0430\u0439\u0442\u043e\u0432 \u0441 \u0437\u0430\u0449\u0438\u0442\u043e\u0439 \u043e\u0442 DDoS, VPS VDS \u0441\u0435\u0440\u0432\u0435\u0440\u044b\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"\ud83e\udd47\u0421\u0442\u0440\u0443\u043a\u0442\u0443\u0440\u0438\u0440\u043e\u0432\u0430\u043d\u0438\u0435 \u0440\u0438\u0441\u043a\u043e\u0432 \u0438 \u0440\u0435\u0448\u0435\u043d\u0438\u0439 \u043f\u0440\u0438 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u0438 BigData \u0434\u043b\u044f \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u0438\u044f \u043e\u0444\u0438\u0446\u0438\u0430\u043b\u044c\u043d\u043e\u0439 \u0441\u0442\u0430\u0442\u0438\u0441\u0442\u0438\u043a\u0438 | ProHoster\" \/>\n\t\t<meta property=\"og:description\" content=\"\u041f\u0440\u0435\u0434\u0438\u0441\u043b\u043e\u0432\u0438\u0435 \u043f\u0435\u0440\u0435\u0432\u043e\u0434\u0447\u0438\u043a\u0430 \u041c\u0430\u0442\u0435\u0440\u0438\u0430\u043b \u0437\u0430\u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043e\u0432\u0430\u043b \u043c\u0435\u043d\u044f, \u0432 \u043f\u0435\u0440\u0432\u0443\u044e \u043e\u0447\u0435\u0440\u0435\u0434\u044c \u0438\u0437-\u0437\u0430 \u0442\u0430\u0431\u043b\u0438\u0446\u044b \u043d\u0438\u0436\u0435:\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki\" \/>\n\t\t<meta property=\"og:image\" content=\"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg\" \/>\n\t\t<meta property=\"og:image:secure_url\" content=\"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg\" \/>\n\t\t<meta property=\"og:image:width\" content=\"350\" \/>\n\t\t<meta property=\"og:image:height\" content=\"350\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2020-03-27T05:42:42+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2020-03-27T05:42:42+00:00\" \/>\n\t\t<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/prohoster\" \/>\n\t\t<meta property=\"article:author\" content=\"https:\/\/www.facebook.com\/prohoster\" \/>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"\ud83e\udd47Structuring risks and solutions when using Big Data to obtain official statistics | ProHoster","description":"Translator's Foreword I was particularly interested in this material because of the table below:","canonical_url":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":null,"og:locale":"en_US","og:site_name":"ProHoster | \u041a\u0443\u043f\u0438\u0442\u044c \u043d\u0430\u0434\u0435\u0436\u043d\u044b\u0439 \u0445\u043e\u0441\u0442\u0438\u043d\u0433 \u0434\u043b\u044f \u0441\u0430\u0439\u0442\u043e\u0432 \u0441 \u0437\u0430\u0449\u0438\u0442\u043e\u0439 \u043e\u0442 DDoS, VPS VDS \u0441\u0435\u0440\u0432\u0435\u0440\u044b","og:type":"article","og:title":"\ud83e\udd47\u0421\u0442\u0440\u0443\u043a\u0442\u0443\u0440\u0438\u0440\u043e\u0432\u0430\u043d\u0438\u0435 \u0440\u0438\u0441\u043a\u043e\u0432 \u0438 \u0440\u0435\u0448\u0435\u043d\u0438\u0439 \u043f\u0440\u0438 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u0438 BigData \u0434\u043b\u044f \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u0438\u044f \u043e\u0444\u0438\u0446\u0438\u0430\u043b\u044c\u043d\u043e\u0439 \u0441\u0442\u0430\u0442\u0438\u0441\u0442\u0438\u043a\u0438 | ProHoster","og:description":"\u041f\u0440\u0435\u0434\u0438\u0441\u043b\u043e\u0432\u0438\u0435 \u043f\u0435\u0440\u0435\u0432\u043e\u0434\u0447\u0438\u043a\u0430 \u041c\u0430\u0442\u0435\u0440\u0438\u0430\u043b \u0437\u0430\u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043e\u0432\u0430\u043b \u043c\u0435\u043d\u044f, \u0432 \u043f\u0435\u0440\u0432\u0443\u044e \u043e\u0447\u0435\u0440\u0435\u0434\u044c \u0438\u0437-\u0437\u0430 \u0442\u0430\u0431\u043b\u0438\u0446\u044b \u043d\u0438\u0436\u0435:","og:url":"https:\/\/prohoster.info\/en\/blog\/administrirovanie\/strukturirovanie-riskov-i-reshenij-pri-ispolzovanii-bigdata-dlya-polucheniya-oficzialnoj-statistiki","og:image":"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg","og:image:secure_url":"https:\/\/prohoster.info\/wp-content\/uploads\/2021\/11\/logo-350.jpg","og:image:width":350,"og:image:height":350,"article:published_time":"2020-03-27T05:42:42+00:00","article:modified_time":"2020-03-27T05:42:42+00:00","article:publisher":"https:\/\/www.facebook.com\/prohoster","article:author":"https:\/\/www.facebook.com\/prohoster"},"aioseo_meta_data":{"post_id":"75636","title":null,"description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"","isEnabled":true},"graphs":[]},"schema_type":null,"schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"seo_analyzer_scan_date":null,"breadcrumb_settings":null,"limit_modified_date":false,"reviewed_by":null,"ai":null,"created":"2021-02-28 17:53:27","updated":"2022-09-28 07:15:59","focus_keyword":null,"additional_keywords":null,"truseo_locale":null},"gt_translate_keys":[{"key":"link","format":"url"}],"_links":{"self":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts\/75636","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/comments?post=75636"}],"version-history":[{"count":0,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/posts\/75636\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/media\/75637"}],"wp:attachment":[{"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/media?parent=75636"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/categories?post=75636"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/prohoster.info\/en\/wp-json\/wp\/v2\/tags?post=75636"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}