• Menu
  • Find a dataset
  • Main content
  • Footer

République
Française

recherche.data.gouv.fr
Access Recherche Data Gouv data repository
    • Recherche Data Gouv at a glance
    • Recherche Data Gouv's organisation
    • Rechercher Data Gouv at the international
    • Join the ecosystem
    • Which research data?
    • Political strategies around data
    • Showcasing my dataset
    • Discover the data platform
    • Recherche Data Gouv Catalogue
    • Geographical (data management clusters and institutional reference centres)
    • According to my discipline (Thematic reference centres)
    • With online resources ("Centres de ressources")
    • on the Recherche Data Gouv repository (thanks to the Platform resource centre)
    • Network of Competence Centres
    • Recherche Data Gouv repository
    • Institutional spaces
    • Trusted repositories
    • Recherche Data Gouv repository guides
    • Recherche Data Gouv catalogue guide
    • Tutorials
    • Legal resources
    • FAQ
  • News
Access Recherche Data Gouv data repository
  1. Home
  2. Guide to entering human sciences and humanities metadata
    • Data accepted into the repository
    • Where to deposit and publish research data
    • Where to deposit Data Management Plans
    • Where to deposit source codes and softwares
    • All you need to know about the Recherche Data Gouv Repository
    • Creating an account
    • Before depositing
    • Depositing a dataset
    • Publishing a dataset
    • Publication process schemas for a dataset
    • Modifying and managing versions of a published dataset
    • Generating a data paper template
    • Withdrawing a published dataset from dissemination
    • Curators' charter
    • The aim of curation
    • Levels of curation
    • In practice
    • The curator's rights regarding datasets
    • Publication process schemas for a dataset
    • Withdrawing a published dataset from dissemination
    • Administrators' charter
    • All you need to know about the Recherche Data Gouv Repository
    • Presentation of a collection
    • Creating a collection
    • Modifying the parameters of a collection
    • Linking a dataset to a collection
    • Complementary features
    • Withdrawing a published dataset from dissemination
    • Browsing through collections
    • Searching for data
    • Displaying and exploring data
    • Guide to entering common metadata
    • Value-lists controled metadata
    • Guide to entering geospatial metadata
    • Guide to entering human sciences and humanities metadata
    • Guide to entering life sciences metadata
    • Guide to entering journal metadata
    • Guide to entering Astronomy and Astrophysics Metadata
    • Guide to entering semantic resource
    • Guide to entering Computational Workflow Metadata
    • Guide to entering file metadata
    • Deposit Cheat-Sheet
    • Ingesting csv files
    • Recommendations on large datasets
    • Curation report template
    • README template
    • DV Uploader
Print

Guide to entering human sciences and humanities metadata

Updated at: 24/07/2026

This guide is currently being compiled. Some of the information may be updated shortly.

 

  1. Social Science and Humanities Metadata
    1. Unit of Analysis
    2. Other Unit of Analysis
    3. Universe
    4. Time Method
    5. Other Time Method
    6. Data collector
    7. Collector Training
    8. Frequency
    9. Sampling Procedure
    10. Other Sampling Procedure
    11. Target Sample Size
    12. Major Deviations for Sample Design
    13. Collection Mode
    14. Other Collection Mode
    15. Type of Research Instrument
    16. Characteristics of Data Collection Situation
    17. Actions to minimize losses
    18. Control Operations
    19. Weighting
    20. Cleaning Operations
    21. Study Level Error Notes
    22. Response Rate
    23. Estimate of Sampling Error
    24. Other Forms of Data Appraisal
    25. Notes
    26. Geographical referential

Social Science and Humanities Metadata

 

Unit of Analysis 

Based on the "Analysis Unit" metadata of the DDI codebook.

Basic unit of analysis or observation that the file describes: individuals, families/households, groups, institutions/organizations, administrative units, and more.

Note: That is a value-list controlled metadata, please see the list below.

  • Individual
  • Organization
  • Family
  • Family: Household family
  • Household
  • Housing Unit
  • Event/Process
  • Geographic Unit
  • Time Unit
  • Text Unit
  • Group
  • Object
  • Other

Individual

 

Other Unit of Analysis 

Other basic unit of analysis or observation that the dataset describes. (Optional)

This variable reports election returns at the constituency level.           

 

Universe 

Based on the "Universe" metadata of the DDI codebook.

Description of the population covered by the data in the file; the group of people or other elements that are the object of the study and to which the study results refer. Age, nationality and residence commonly help to delineate a given universe, but any number of other factors may be used, such as age limits, sex, marital status, ethnic group, nationality, income, and more. The universe may consist of elements other than persons, such as housing units, court cases, deaths, countries, and so on. In general, it should be possible to tell from the description of the universe whether a given individual or element is a member of the population under study. Also known as the universe of interest, and target population. (Optional)

Click on the (+) button to add an universe.

Population residing in metropolitan France (excluding Corsica): between 18 and 79 years old; -in main residence; -reading enough French to answer a self-administrated questionnaire.

 

Time Method 

Based on the "Time Method" metadata of the DDI codebook.

The time method or time dimension of the data collection, such as panel, cross-sectional, trend, time-series, or other. (Optional)

Note: That is a value-list controlled metadata, please see the list below.

  • Longitudinal
  • Longitudinal: Cohort/Event-based
  • Longitudinal: Trend/Repeated cross-section
  • Longitudinal: Panel
  • Longitudinal: Panel: Continuous
  • Longitudinal: Panel: Interval
  • Time Series
  • TimeSeries: Continuous
  • TimeSeries: Discrete
  • Cross-section
  • Cross-section ad-hoc follow-up
  • Other

Time Series

 

Other Time Method 

Other time method or time dimension of the data collection. (Optional)

Panel survey

 

Data collector 

Based on the "Data Collector" metadata of the DDI codebook.

Individual, agency or organization responsible for administering the questionnaire or interview or compiling the data. (Optional)

Questionnaire Administration, Survey Research Center (SRC), University of Michigan.

 

Collector Training 

Based on the "Collector Training" metadata of the DDI codebook.

Type of training provided to the data collector. (Optional)

Describe the research project, describe the population and sample, suggest methods and language for approaching the subjects, explain questions and key terms of the survey instrument.

 

Frequency 

Based on the "Frequency of Data Collection" metadata of the DDI codebook.

If the data collected includes more than one point in time, indicate the frequency with which the data was collected; That is, monthly, quarterly, or other. (Optional)

Monthly

 

Sampling Procedure 

Based on the "Sampling Procedure" metadata of the DDI the codebook.

Type of sample and sample design used to select the survey respondents to represent the population. May include reference to the target sample size and the sampling fraction.

Note: That is a value-list controlled metadata, please see the list below.

  • Total universe/Complete enumeration
  • Probability
  • Probability: Simple random
  • Probability: Systematic random
  • Probability: Stratified
  • Probability: Stratified: Proportional
  • Probability: Stratified: Disproportional
  • Probability: Cluster
  • Probability: Cluster: Simple random
  • Probability: Cluster: Stratified random
  • Probability: Multistage
  • Non-probability
  • Non-probability: Availability
  • Non-probability: Purposive
  • Non-probability: Quota
  • Non-probability: Respondent-assisted
  • Mixed probability and non-probability

Total universe/Complete enumeration

 

Other Sampling Procedure 

Other type of sample or sample design used to select the survey respondents. 

Samples sufficient to produce approximately 2,000 families with completed interviews were drawn in each state. Families containing one or more Medicaid or uninsured persons were oversampled.

 

Target Sample Size 

Based on the "Target Sample Size" metadata of the DDI codebook.

Specific information regarding the target sample size, actual sample size, and the formula used to determine this.

  • Actual: Specific information regarding the target sample size, compared to its actual size. (Optional)
  • Formula: Specific information regarding the target sample size, compared to the formula used to determine its size. (Optional)

  • Actual: 385
  • Formula: n0=Z2pq/e2=(1.96)2(.5)(.5)/(.05)2=385 individuals

 

Major Deviations for Sample Design 

Based on the "Major Deviations from the Sample Design" metadata of the DDI codebook.

Show correspondence as well as discrepancies between the sampled units (obtained) and available statistics for the population (age, sex-ratio, marital status, etc.) as a whole. (Optional)

Before adjustment, following categories appear under-estimated in the sample of survey respondents :- Single person households; - Paris basin housings; - Households with individuals over 60 years old, or without individual under 25 years old; - Under graduated persons.

 

Collection Mode 

Based on the "Mode of Data Collection" metadata of the DDI codebook.

Method used to collect the data; instrumentation characteristics (e.g., telephone interview, mail, questionnaire, or other). (Optional)

Note: That is a value-list controlled metadata, please see the list below.

  • Interview
  • Face-to-face interview
  • Face-to-face interview: CAPI
  • Face-to-face interview: PAPI
  • Telephone interview
  • Telephone interview: CATI
  • E-mail interview
  • Web-based interview
  • Self-administered questionnaire
  • Fixed form self-administered questionnaire
  • Fixed form self-administered questionnaire: E-mail
  • Fixed form self-administered questionnaire: Paper
  • Fixed form self-administered questionnaire: SMS/MMS
  • Fixed form self-administered questionnaire: Web-based
  • Interactive self-administered questionnaire
  • Interactive self-administered questionnaire: CASI
  • Interactive self-administered questionnaire: CASI: VCASI
  • Interactive self-administered questionnaire: CASI: ACASI
  • Interactive self-administered questionnaire: CASI: TACASI
  • Interactive self-administered questionnaire: CAWI
  • Focus group
  • Face-to-face focus group
  • Telephone focus group
  • Online focus group
  • Self-administered writings and/or diaries
  • Self-administered writings and/or diaries: E-mail
  • Self-administered writings and/or diaries: Paper
  • Self-administered writings and/or diaries: Web-based
  • Observation
  • Field observation
  • Participant field observation
  • Non-participant field observation
  • Laboratory observation
  • Participant laboratory observation
  • Non-participant laboratory observation
  • Computer-based observation
  • Experiment
  • Laboratory experiment
  • Field/Intervention experiment
  • Web-based experiment
  • Recording
  • Content coding
  • Transcription
  • Compilation/Synthesis
  • Summary
  • Aggregation
  • Simulation
  • Measurements and tests
  • Educational measurements and tests
  • Physical measurements and tests
  • Psychological measurements and tests
  • Other

Telephone interview

 

Other Collection Mode 

Other method used to collect the data. (Optional)

Mail questionnaires

 

Type of Research Instrument 

Based on the "Type of Research Instrument" metadata of the DDI codebook.

Type of data collection instrument used. “Structured” indicates an instrument in which all respondents are asked the same questions/tests, possibly with pre coded answers. If a small portion of such a questionnaire includes open-ended questions, provide appropriate comments. “Semi-structured” indicates that the research instrument contains mainly open-ended questions. “Unstructured” indicates that in-depth interviews were conducted. (Optional)

Structured

 

Characteristics of Data Collection Situation 

Based on the "Characteristics of Data Collection Situation" metadata of the DDI codebook.

Description of noteworthy aspects of the data collection situation. Includes information on factors such as cooperativeness of respondents, duration of interviews, number of call backs, or similar. (Optional)

There were 1,194 respondents who answered questions in face-to-face interviews lasting approximately 75 minutes each.

 

Actions to minimize losses 

Based on the "Actions to Minimize Losses" metadata of the DDI codebook.

Summary of actions taken to minimize data loss. Include information on actions such as follow-up visits, supervisory checks, historical matching, estimation, and so on. (Optional)

To minimize the number of unresolved cases and reduce the potential nonresponse bias, four follow-up contacts were made with agencies that had not responded by various stages of the data collection process.

 

Control Operations 

Based on the "Control Operations" metadata of the DDI codebook.

Control operations methods to facilitate data control performed by the primary investigator or by the data archive. (Optional)

Ten percent of data entry forms were reentered to check for accuracy.

 

Weighting 

Based on the "Weighting" metadata of the DDI codebook.

The use of sampling procedures might make it necessary to apply weights to produce accurate statistical results. Describes the criteria for using weights in analysis of a collection. If a weighting formula or coefficient was developed, the formula is provided, its elements are defined, and it is indicated how the formula was applied to the data. (Optional)

The pilot sample selection procedure underestimated the impact that the small sample size and the primary unit selection process could have on the dispersion of the weights. To limit this dispersion, it was proposed to set a standardized sampling weight for the entire sample. Individual weights are adjusted for non-response using the homogeneous response group method, then by calibration on five margins of the “Enquête annuelle du recensement 2014” (sex, age, nationality, diploma, ZEAT). The weights of only the survey respondents are adjusted again to the same margins to correct for non-response. Their sum is equal to the sample size of respondents. More details about the adjustment procedure and the use of weightings in the weighting documentation.

 

Cleaning Operations 

Based on the "Weighting" metadata of the DDI codebook.

Method used to clean the data collection, such as consistency  checking, wildcode checking, or other. (Optional)

Checks for undocumented codes were performed, and data were subsequently revised in consultation with the principal investigator.

 

Study Level Error Notes 

Note elements used for any information annotating or clarifying the methodology and processing of the study. (Optional)

 

 

Response Rate 

Based on the "Response Rate" metadata of the DDI codebook.

Percentage of sample members who provided information. (Optional)

Response rate: 84.9% of 2.770 asked to respond the survey.

 

Estimate of Sampling Error 

Based on the "Estimates of Sampling Error" metadata of the DDI codebook.

Measure of how precisely one can estimate a population value from a given sample. (Optional)

To assist NES analysts, the PC SUDAAN program was used to compute sampling errors for a wide-ranging example set of proportions estimated from the 1996 NES Pre-election Survey dataset. For each estimate, sampling errors were computed for the total sample and for twenty demographic and political affiliation subclasses of the 1996 NES Pre-election Survey sample. The results of these sampling error computations were then summarized and translated into the general usage sampling error table provided in Table 11. The mean value of deft, the square root of the design effect, was found to be 1.346. The design effect was primarily due to weighting effects (Kish, 1965) and did not vary significantly by subclass size. Therefore the generalized variance table is produced by multiplying the simple random sampling standard error for each proportion and sample size by the average deft for the set of sampling error computations.

 

Other Forms of Data Appraisal 

Based on the "Other Forms of Data Appraisal" metadata of the DDI codebook.

Other issues pertaining to the data appraisal. Describe issues such as response variance, non-response rate and testing for bias, interviewer and response bias, confidence levels, question bias, or similar. (Optional)

These data files were obtained from the United States House of Representatives, who received them from the Census Bureau accompanied by the following caveats: "The numbers contained herein are not official 1990 decennial Census counts. The numbers represent estimates of the population based on a statistical adjustment method applied to the official 1990 Census figures using a sample survey intended to measure overcount or undercount in the Census results. On July 15, 1991, the Secretary of Commerce decided not to adjust the official 1990 decennial Census counts (see 56 Fed. Reg. 33582, July 22, 1991). In reaching his decision, the Secretary determined that there was not sufficient evidence that the adjustment method accurately distributed the population across and within states. The numbers contained in these tapes, which had to be produced prior to the Secretary's decision, are now known to be biased. Moreover, the tapes do not satisfy standards for the publication of Federal statistics, as established in Statistical Policy Directive No. 2, 1978, Office of Federal Statistical Policy and Standards. Accordingly, the Department of Commerce deems that these numbers cannot be used for any purpose that legally requires use of data from the decennial Census and assumes no responsibility for the accuracy of the data for any purpose whatsoever. The Department will provide no assistance in interpretation or use of these numbers."

 

Notes 

Based on the "Notes and comments" metadata of the DDI codebook.

General notes about this Dataset.

  • Type: Type of note. (Optional)
  • Subject: Note subject. (Optional)
  • Text: Text for this note. -This field will become required if you choose to enter values in one or more of this optional fields.

Example #1:

  • Type: Description
  • Subject: Data
  • Text: The variables in this study are identical to earlier waves.

Example #2:

  • Type: Information
  • Subject: Files
  • Text: Data on employment and income refer to the preceding year, although demographic data refer to the time of the survey.

 

Geographical referential 

Geographical referential used in the Dataset

  • Level: Level of the geographical referential of the Dataset
  • Version: Version of the geographical referential of the Dataset. (Optional)

Click on the (+) button to add a geographical referential.

 

 

Access Recherche Data Gouv data repository

Ministère
de l'Enseignement
supérieur,
de la Recherche
et de l'Espace

Contact Recherche Data Gouv
Access the contact form
Talk about Recherche Data Gouv
Access to the communication kit

Follow us
on social networks

  • legifrance.gouv.fr
  • gouvernement.fr
  • service-public.fr
  • data.gouv.fr
  • Legal notices
  • Releases notes
  • Sitemap
  • Accessibility: non-compliant
  • Cookies management

Unless otherwise stated, all content on this site is under licence etalab-2.0, source code is under license GNU GPL V3.

Back to top