Start here
Standalone self-paced course

SurveyCTO for field systems

Design a form, explain its calculations, collect and review submissions, then connect two forms through a server dataset and a case list. The course uses invented household records throughout. A second practical uses the supplied roster and education questionnaire and ten-case household list.

27 modulesFoundations through advanced workflowsCalculations, datasets, and casesKnowledge checks and labsCourse file works offline
Course author

Aubrey Jolex

Senior Research Associate
Estimated commitment

14โ€“18 hours for lessons and checks
16โ€“24 hours for form and server labs
6โ€“10 hours for the first capstone
8โ€“12 hours for the questionnaire practical

Plan about eight weeks at six to eight hours per week, or use the course as a reference while building a practice project.

How to study

Begin with Module 1's illustrated, first-time walkthrough. It shows where to sign in, how a tiny form moves from Design to collection, and how to verify a server submission and CSV export. From Module 2 onward, use the Practice PDF in the top bar and complete each Build the questionnaire as you learn milestone. Keep that PDF in a second tab beside the lesson; by Modules 25โ€“27 you will be integrating and release-testing work already started, not meeting the questionnaire for the first time. Each module also has a knowledge check. It allows three complete attempts; after the third, the answer key opens the next module.

Your answers and module progress are stored in this browser. The course HTML embeds its practice workbooks, CSVs, and questionnaire PDF. A SurveyCTO server, the product documentation, and collection tests require a connection. The course is not a server account, and completing a simulated exercise does not publish real data.

Two complementary practice tracks

Build-along questionnaire: use the ten-household roster and education Practice PDF for short milestones in form structure, logic, repeats, case selection, collection, and exports. Connected-data example: use HH001โ€“HH004 to learn server datasets, publishing, and a follow-up form. These are different synthetic cases; do not merge their files or IDs.

Modules 2โ€“5First fields, consent and branches
โ†’
Modules 8โ€“11Ages, roster and device cycle
โ†’
Modules 19โ€“22Cases and export checks
โ†’
Modules 25โ€“27Integration and release test

The running example

A field team lists a small set of households. The listing form records a stable household ID, head name, consent, household members, and an adult count. A server dataset stores the current household lookup. A follow-up form uses the same ID to pull the household's name and contact state. A cases dataset assigns the follow-up to a collector. Each handoff has a specific test: submission received, dataset row updated, attachment refreshed, case visible, and follow-up submitted.

Listing form
one submission
publishes โ†’
Household dataset
one row per ID
attaches โ†’
Follow-up form
preloaded values
guided by โ†’
Cases dataset
collector queue

The same architecture applies to schools, clinics, plots, and other tracked subjects. The example uses no real respondent information.

Files and ground rules

  • Download the practice XLSForms and CSVs using the buttons below. They contain only synthetic values.
  • Use a trial or designated training server for live labs. Keep its URL and server admin credentials outside the workbook.
  • Never upload a real sample list, identifiable data, or internal protocol to the practice server.
  • Keep a test log with input, expected behavior, observed behavior, form version, and device.

The first capstone uses HH001โ€“HH004. The later questionnaire practical uses HH-2024-001 through HH-2024-010 and its own files. The completed files are examples to inspect after attempting the labs. Their form IDs and dataset names are for training; verify every field and publishing link before deploying them on a server.

Questionnaire practical: integration and release

The build-along milestones begin in Module 2. Modules 25โ€“27 bring those pieces together with the supplied ten-household list, compare your form with a reference, and run the twelve-test matrix. The reference exposes decisions the paper source does not settle; it is not an approved production instrument.

Completion evidence

Keep the edited form source, a calculation test matrix, the two-form mapping, a cases CSV, synthetic submissions, a downloaded dataset CSV, a role test, and an incident/release note. A course percentage shows that you opened the checks; these artifacts show what the field system actually does.

Module 1 ยท Part I ยท Foundations

The SurveyCTO system

From your first login to a verified CSV export

By the end of this module, you can:
  • Explain the six-stage design-to-export workflow
  • Set up a safe practice route and identify the server areas
  • Prove that a test interview reached the server and CSV export

First, know what you are operating

SurveyCTO is not just a form editor. It has a server that holds form versions, users, incoming interviews, and exports; SurveyCTO Collect, the phone or tablet app used by many field teams; and SurveyCTO Desktop, a computer application for synchronizing and exporting data. A respondent can also complete an enabled web form in a browser. You do not need all of these components for every study, but you must know which one holds the form, which one holds an unfinished interview, and when data actually reach the server.

Your goal in this module: follow one tiny practice form from creation to a CSV export. You do not need to understand datasets, case management, calculations, or publishing links yet. Those come after you have seen the basic route.
ComponentWhat a beginner does thereWhat it does not prove by itself
Server consoleDesign, deploy, manage access, monitor submissions, and export data.A form on the server is not automatically on every device.
SurveyCTO CollectDownload blank forms, fill them offline, finalize, and send when connected.A finalized interview on the device is not yet a server submission.
Web formFill an enabled form in a browser, useful for a first end-to-end test.A form preview is not necessarily a retained submission.
SurveyCTO DesktopSynchronize and export to a local folder, including CSV files for analysis.It is not the place where you design questions.

Read the whole workflow before opening a menu

Six-step SurveyCTO cycle: design a form, publish it to the server, download it to a device, collect an interview, sync it to the server, then review and export the data.
One common mobile workflow. The supplied illustration gives the right order. In the current console, the form-release action is commonly called Save and deploy. That is different from the later, optional feature called Publish into, which copies selected submission values into a server dataset. A web form can skip the device-download step, but still needs a deployed, accessible form.
  1. DesignWrite the questions and answer rules in the online designer or a SurveyCTO XLSForm. Preview a draft with fake answers. Evidence: the intended question and options appear in Test.
  2. DeploySave and deploy a tested form to the practice server. This is the version collectors can use. Evidence: the deployed form appears under Design with the intended form ID and version.
  3. DownloadOn a phone or tablet, sign in to the right server and use Get Blank Form. An enabled web form instead opens from Collect. Evidence: the form is available to fill on that device or browser.
  4. CollectFill the form, check answers, and finalize it. Work may happen without mobile Internet. Evidence: the interview is saved or finalized locally; this alone does not confirm server receipt.
  5. SendWhen connected, use Send Finalized Form in Collect (or submit the web form). Evidence: a successful send and the corresponding submission visible on the server.
  6. Review and exportUse Monitor to inspect incoming interviews and Export to download a CSV or another supported analysis format. Evidence: the expected test ID and answers appear in the exported file.

The image loops back to Design because field testing often reveals a wording or logic change. A revised form has to be tested, deployed, and received by devices again. Do not mistake โ€œI edited the draftโ€ for โ€œeveryone now has the update.โ€

Set up a safe practice space

  1. Choose a server you may use. Use an approved training server or your own trial account. Record its server name/address in your private notes; never put credentials in a workbook or share them in the course. Do not experiment on a live project server.
  2. Confirm your role. A designer or administrator can create and deploy forms. A data-collection account can fill permitted forms. Review and export access may be restricted. If a control below is missing, ask the server owner which role you have; do not try to bypass permissions.
  3. Use synthetic data. The practice files use invented IDs such as HH001. Do not upload a real sample list, respondent name, production API key, or confidential questionnaire to a trial server.
  4. Pick one collection route. For the quickest first run, use web data collection if it is enabled for your practice server. For the field route, install SurveyCTO Collect on a test device and configure it with the practice server name and a permitted collector account.
  5. Keep one test log. Note form ID, version, test ID, route (web or device), expected answer, actual answer, and whether the server received it. This becomes your evidence when something fails.
No server access yet? You can still complete the map, inspect the supplied starter XLSForm in Module 2, and plan the test route. Mark the live steps below as โ€œnot run,โ€ not โ€œpassed.โ€ Ask the owner for a training account before making claims about deployment, collection, or export.

Find your way around the console

Sign in to the practice server and locate these five areas. Menu labels can vary slightly by permissions and interface version, so match the purpose rather than guessing from position alone.

AreaFirst thing to locateQuestion it answers
DesignYour forms and datasets; form designer; Test; deployed version.What will a collector be asked?
CollectWeb data collection and the Fill out action, if enabled.How can this form be filled now?
MonitorForm submissions and, when available, Data Explorer.Did the server receive the test interview?
ExportThe form under Your data and its download options.Can the test answers be taken into a CSV for analysis?
ConfigureYour users and their roles; inspect only unless you are authorized to change them.Who may design, collect, review, or export?

The server connects these areas; the tabs are not separate databases. SurveyCTO's more advanced server datasets can provide preloaded values, current records, and case lists. They are not required for this first test. A CSV is a file you can upload or export; a server dataset is a managed object that can be downloaded as a CSV. Module 12 starts the connected-data pathway.

Make one complete trip with a tiny form

These steps deliberately use the online designer so a first-time learner can see the complete cycle before studying the spreadsheet source in Module 2. Use a unique practice ID such as first_trip_yourinitials; do not reuse an ID already on the server.

  1. Design: In Design โ†’ Your forms and datasets, choose + โ†’ Start new form. Give it the visible title โ€œMy first practice formโ€ and a stable, space-free form ID. Choose Edit online. Add a Text field with the question โ€œWhat is your practice code?โ€ and field name practice_code. Save the field.
  2. Test: Switch the designer to Test. Enter HH001. Confirm the label and input work. If the question is missing, return to Design and check the saved field; a successful editor load alone is not a test.
  3. Deploy: Use Save and deploy, then return to the server's Design area. Note the form ID and deployed version. Testing a draft and deploying it are separate actions.
  4. Collect by web, if enabled: In Collect, open Fill out for this form. Enter HH001, mark the form finalized, and submit it. If web collection is off or unavailable to your role, use the device route instead.
  5. Collect by device, alternative route: In SurveyCTO Collect, configure the practice server and collector credentials. Choose Get Blank Form, select this form, then Fill Blank Form. Enter HH001 and finalize. When connected, choose Send Finalized Form. Saving locally without sending is not the final step.
  6. Verify receipt: In Monitor, refresh the form's submission count and inspect the new submission. Match the practice code and submission time. If there is no record, first check whether you used the correct server, deployed form, account, and send/submit step.
  7. Export: In Export, choose the same form and download its data as CSV. Open the file and locate the practice_code column with HH001. Keep the export with your test log. The export is a snapshot, not an automatically updating link to later interviews.
What success looks like: one form ID under Design, one sent HH001 interview under Monitor, and one CSV row containing HH001. The field label, stored field name, server submission, and exported CSV column should line up. If you ran both web and device routes, you may have two submissions; record both rather than treating one as an accidental duplicate.

Five words that must not blur together

TermMeaning in the HH001 exampleCommon mistake
Form definitionThe questions, choices, logic, form ID, and version you designed.Thinking the blank form is already an interview.
Draft / deployed versionA draft is being tested; a deployed version can be received by collectors.Assuming every saved edit reached phones.
Saved / finalized / sentStages of a device interview before and during upload.Assuming โ€œfinalizedโ€ means the server received it.
SubmissionOne interview received by the server, with its own submission KEY.Using the respondent name as the interview's unique ID.
CSV exportA downloaded file representing the selected data at export time.Assuming an old CSV includes new submissions automatically.

Later, the course adds server datasets, which maintain rows that can be preloaded into forms or updated from submissions, and cases datasets, which drive a collector's case list. A form ID names the questionnaire; a household ID such as HH003 names the household; a submission KEY names one particular interview. Those IDs may appear together but are not interchangeable. Module 14 returns to dataset keys, and Module 19 to case IDs.

Practice 1 ยท first end-to-end cycle

Prove each handoff

  1. Point to each of the six numbered stages in the image and name the SurveyCTO place or device where it happens.
  2. Complete the tiny practice_code exercise on a practice server, or write โ€œnot runโ€”no server accessโ€ beside the live steps.
  3. For the test code HH001, record the form ID and version, collection route, local completion state, server submission evidence, and exported CSV filename.
  4. Explain in one sentence why Save and deploy, Send Finalized Form, and Publish into a dataset are three different actions.
Model answer and troubleshooting cues

Design and deploy release a form version; Send Finalized Form uploads an interview from a device; Publish into is an optional server rule that copies selected submission values into a managed dataset. If HH001 is absent in Monitor, check the account/server, whether the correct version was deployed and downloaded, and whether the interview was actually sent. If it is in Monitor but absent from your CSV, check the selected form, export filters, and export time before changing the questionnaire.

Knowledge check

Module 1 check

Attempts remaining: 3
1. Which sequence describes the basic mobile route?
2. A collector finalized HH001 offline. What has definitely happened?
3. Where should you first verify that the server received the interview?
4. What does Save and deploy release?
5. What is a CSV export?
Module 2 ยท Part I ยท Foundations

Read an XLSForm

The survey, choices, and settings sheets

By the end of this module, you can:
  • Find a field and its choices
  • Use SurveyCTO column names
  • Keep stable form and variable IDs
Where this fits after Module 1: The online designer let you make a small practice form without opening Excel. An XLSForm is another editable definition of a SurveyCTO formโ€”not an export of submitted answers. The survey, choices, and settings worksheets describe what the collector will see and what values SurveyCTO will store.

Start from the maintained template

A SurveyCTO XLSForm is an Excel workbook with multiple worksheets, not a CSV. The official template includes survey, choices, and settings plus help and audit material. Copy it for a new form. Do not rebuild the three sheets from an empty workbook; that discards useful guidance and may change defaults. When editing an existing form, start from its maintained source and preserve its form ID, variable names, and version history unless a documented change requires otherwise.

SheetQuestion to askExample
surveyWhat is asked, calculated, shown, or repeated?text | hhid | Household ID
choicesWhat codes and labels are offered?yesno | yes | Yes
settingsWhat is this form called and identified as?form_id = hh_listing_training

Follow a question through the workbook

survey:  type=select_one yesno   name=consent   label=May we continue?
choices: list_name=yesno value=yes label=Yes
choices: list_name=yesno value=no  label=No
settings: form_id=hh_listing_training

The word after select_one must match choices.list_name. The submitted value is yes or no, not the English label. Logic such as ${consent} = 'yes' compares the stored code. The name becomes a data column, so choose it carefully and keep it stable across a live form's versions.

SurveyCTO differs from other XLSForm platforms in small but important ways. Use relevance for skip logic and value for a choice code. Use the exact headers constraint message and required message with spaces. These are not interchangeable with another platform's column names.

Version the source as well as the server

Keep one maintained workbook for each form. Before a change, copy it to a dated or versioned working file, record the purpose of the change, and preserve the prior deployed version. A different filename does not itself create a new form: the form ID in settings is the stable identity. An accidental form-ID change can break case rows, attachments, publishing links, and device update expectations.

Source record

hh_listing_training_v20260925.xlsx contains form_id=hh_listing_training. Its change log says โ€œAdded an adult count and consent filter; tested consent No, three-member household, and changed ages.โ€ The version in the file and the version uploaded to the server should be traceable to the same test note.

Practice 2 ยท workbook reading

Find the stored answer and the stable ID

  1. Download the listing starter XLSForm from Start here.
  2. Locate the hhid and consent rows in survey.
  3. Find the yesno list in choices. Write down the label and value for each option.
  4. Read settings and note the form ID. Save a working copy without changing that ID.
Expected result

The consent question is select_one yesno. Its response is a stored code from choices.value. The training form ID is stable even if you rename your working file.

Knowledge check

Module 2 check

Attempts remaining: 3
1. Which worksheet defines questions and calculations?
2. Which choice column stores the answer code?
3. Which identifier should stay stable across versions of one form?
Module 3 ยท Part I ยท Foundations

Choose field types and codes

Questions that produce usable data

By the end of this module, you can:
  • Select a type from the answer
  • Separate labels from stored values
  • Write clear hints and messages

Choose a type from the answer

A variable that looks numeric may still be an identifier. HH003 is text. A phone number is text because leading zeros, a plus sign, and formatting matter; no arithmetic should be performed on it. A count of people is an integer. An area in hectares may be a decimal. A single status is select_one; several facilities are select_multiple. A note gives guidance without collecting an answer. A calculate stores a derived value without asking the user.

PromptTypeStored exampleRisk if mistyped
Household IDtextHH003Numeric input cannot preserve the ID
Member age in completed yearsinteger17Free text defeats numerical checks
Distance in kmdecimal1.75Integer rejects valid fractions
Visit outcomeselect_onecompleteFree text creates spelling variants
Facilities usedselect_multiplewater toiletSelect One loses combinations

Keep labels and codes separate

Write the wording that the collector or respondent sees in label. Write short, stable machine values in choices.value. A translated form can show another language while preserving the same stored code. A changed label is a documentation and translation task; a changed code can break comparisons, publishing filters, and analysis across versions.

Make the question actionable. โ€œHousehold size?โ€ leaves uncertainty about visitors and absent members. โ€œHow many people usually live in this household?โ€ followed by a project-approved definition gives a consistent count. Use hint for a brief field instruction, and a constraint message that says how to correct an invalid answer. Avoid forcing an answer when the respondent may not know it; choose a missingness policy and code it deliberately.

Design for analysis before programming

For every field in the practice instrument, make a one-line specification with question name, type, label, allowed values, missingness rule, and intended use. This prevents a later calculation from assuming an input is always available when the questionnaire allows it to be skipped. A field for phone number might be optional, while hhid should be required and checked against the sample or case list.

Field specification

member_age: integer, โ€œHow old is this member in completed years?โ€, 0โ€“120, required for rostered members. The adult-count calculation uses member_age >= 18. Boundary tests are 17, 18, 120, and 121.

Practice 3 ยท field choices

Specify five fields

  1. Write types for household ID, member age, phone number, visit outcome, and facilities used.
  2. Choose stable values for visit outcome: complete, callback, and refused.
  3. Write one hint that prevents a foreseeable collector error and one constraint message that tells the collector how to fix a failed age.
  4. Explain which of these fields might feed a publishing filter and why its code must remain stable.
Model answer

Household ID and phone number are text; age is integer; outcome is Select One; facilities are Select Multiple. A status code such as complete can feed a calculated 1/0 publishing filter, so its stored value must be stable.

Knowledge check

Module 3 check

Attempts remaining: 3
1. An ID may begin with 00. Which type preserves it?
2. A question allows several facilities. Which type fits?
Module 4 ยท Part I ยท Foundations

Preview and deploy

Draft tests, live tests, and versions

By the end of this module, you can:
  • Test a draft
  • Deploy a practice form
  • Record what changed
From a first trip to a release test: Module 1 proved that one simple form could travel from Design to a CSV export. It did not prove consent branches, age boundaries, repeat behavior, or device updates. This module turns that initial smoke test into a documented release check.

Run a test that can fail

Preview is a draft-form test. It is useful only when you enter cases that challenge the logic. For the listing form, test consent No; consent Yes with one member; age 17 and 18; a blank required ID; invalid age 121; and a change from Yes to No after answering follow-ups. Record both expected and observed behavior. A form that opens successfully has passed only a load test.

CaseInputExpected
Refusalconsent=noMember questions are hidden and not required
BoundaryAge 17, then 18Adult count changes from 0 to 1
InvalidAge 121Correction message prevents proceeding
CorrectionConsent Yes, roster data, then NoInspect hidden answers and final stored values

Deploy, then test the full path

  1. Save the maintained XLSForm and upload it as a draft. Read any validation message and correct the named row or expression.
  2. Preview the test matrix. Keep a record of the workbook version and date.
  3. Deploy the practice form only after the critical paths pass.
  4. Open the deployed form through the method collectors will use. Submit one synthetic record.
  5. Find that record in the server's submission list and inspect its exported values. If a publishing link exists, verify the dataset row separately.

The first live record tests access, device or browser delivery, finalization, sending, server receipt, and export. It does not prove every questionnaire branch works; that is why the earlier matrix remains necessary.

Protect existing records during updates

Once collection starts, a change to a field name or stored choice value changes the meaning of old and new data. A new required field can make editing an older submission difficult. A changed relevance rule can alter what an edited record retains. Test an update on a copy with synthetic submissions before updating a live project. Tell field devices how to receive the new form. Keep old source versions and a release note.

Release note for the practice form

Form ID hh_listing_training; version 20260925; changed adult-count calculation; tests: 17, 18, three-member roster, edited age; device update: obtain new form while connected; outstanding issue: iOS test not run. The note says what changed and what remains uncertain.

Practice 4 ยท release test

Write a pass/fail record

  1. Create a five-row test matrix using the cases above.
  2. Run it against your draft. Record exact inputs, result, and form version.
  3. If you have a training server, deploy and send one invented household. Find its submission and note the KEY.
  4. Write a two-sentence release message telling collectors how to obtain the version and which previous records need caution.
Completion rule

A screenshot of the preview is not sufficient. Keep the test inputs and a visible server receipt for the live submission.

Knowledge check

Module 4 check

Attempts remaining: 3
1. What does preview establish?
2. After updating a deployed form, what must field devices do?
Module 5 ยท Part II ยท Calculations and logic

Relevance, required, and constraints

Three different rules for field behavior

By the end of this module, you can:
  • Write a consent branch
  • Test a numerical boundary
  • Check blanks and changed answers

Three rules answer three different questions

PropertyQuestionSchool-visit example
relevanceShould the field appear now?Show details when ${consent} = 'yes'
requiredMay the user leave a visible field blank?Require household ID
constraintIs an entered value acceptable?Allow age from 0 to 120

These rules do not replace each other. An optional age with a range constraint can still be blank. A required field hidden by relevance should not block a refusal path. A field may be visible and valid but still not have the meaning your question intended; wording and training matter too.

Write and test a consent branch

Write the rule in words first: โ€œAsk the roster only if consent is Yes.โ€ Then write it as ${consent} = 'yes' in the relevance column on the begin group row. The group rule applies to the member questions within it. Use a stored choice value, not the visible label. If you have a refusal follow-up, put it outside the group and give it the opposite relevance.

type               name           label                   relevance
select_one yesno   consent        May we continue?
begin group        interview      Household interview     ${consent} = 'yes'
integer            hh_size        Usual residents
end group          interview
text               refusal_note   Reason, if offered      ${consent} = 'no'

Preview three paths: consent unanswered; Yes; No. Then choose Yes, enter a value, return to consent and change to No. Inspect what the final record stores, not only what disappears from view.

Test a hard range and a multiple-choice condition

For a member age use . >= 0 and . <= 120 with a message such as โ€œEnter an age from 0 to 120, or follow the missing-age procedure.โ€ Test โˆ’1, 0, 120, and 121. A constraint is a hard stop; if unusual but possible ages need supervisor review, design that as a separate soft-check process rather than widening or disabling the hard rule without thought.

For select_multiple facilities, the answer may hold several codes such as water toilet. To show an Other text box use selected(${facilities}, 'other'). A direct equality test against 'other' will fail for a combined answer. Test Other alone and Other with another selection, then remove Other and see whether a stale explanation remains in the saved result.

Expression drill

Write the relevance expression for โ€œSpecify other facilityโ€ when Other is selected in facilities.

Practice 5 ยท consent and validation

Build both branches

  1. Add a consent question, a conditional group, and a refusal note to your working XLSForm.
  2. Mark household ID required. Give age a hard range and correction message.
  3. Test unanswered, Yes, No, Yesโ†’No, age boundaries, Other alone, and Other with water.
  4. Record the final stored values for a refusal. If the observed data differs from the intended rule, revise the form before moving on.
Expected interpretation

Relevance controls visibility, required controls a visible blank, and constraint checks an entered value. The Select Multiple test must use selected().

Knowledge check

Module 5 check

Attempts remaining: 3
1. Which column controls whether a field appears?
2. Does a constraint alone force an optional blank answer?
3. For a multiple-choice field storing several codes, which test is appropriate?
Module 6 ยท Part II ยท Calculations and logic

Calculate fields: the core model

Inputs, outputs, types, and recomputation

By the end of this module, you can:
  • Write an if() calculation
  • Trace dependency changes
  • Choose calculate or calculate_here

A calculate field stores an answer nobody types

A calculate row has a name and an expression in calculation. It is hidden from the collector but appears in submitted data. Use it to create a derived value that has a defined meaning, such as is_adult, adult_count, or eligible_for_followup. Do not add a calculated field only because a formula can be written. Specify its inputs, output type, possible blank value, when it should change, and who will use it.

type        name         calculation
calculate   is_adult     if(${member_age} >= 18, 1, 0)
calculate   visit_flag   if(${visit_outcome} = 'complete', 1, 0)

These values can drive relevance, checks, publishing filters, or analysis. The displayed label on visit_outcome can change; the stored code complete must remain consistent with the expression.

Track dependencies and recalculation

Suppose a collector enters age 17, then corrects it to 18. A regular calculate field depending on age recalculates, changing is_adult from 0 to 1. Any later total that depends on is_adult must also update. Build a dependency chain on paper before programming:

member_age
entered
feeds โ†’
is_adult
0 or 1
aggregated โ†’
adult_count
household total
used by โ†’
followup_flag
routing

A test that enters only the final value misses stale-calculation bugs. Test the correction. Also test whether a field becomes irrelevant: a calculation placed inside an irrelevant group may be blank rather than 0. Those states must have distinct interpretations.

Choose calculate_here rarely and deliberately

SurveyCTO's regular calculate fields recalculate as needed when forms load, save, or dependencies change. calculate_here runs when the user reaches that location in the form. It fits a checkpoint such as โ€œtime first reached the consent screen,โ€ but is usually the wrong tool for a count that should react to edits. A one-time timestamp can use once(format-date-time(now(), '%Y-%b-%e %H:%M:%S')) in a calculate_here row. Test leaving and returning to the screen and editing a saved record.

Do not use once() to hide a broken dependency in a total. It freezes a value that should probably change when source answers are corrected.

Calculation drill

Write a numeric 1/0 calculate expression for is_adult from member_age.

Practice 6 ยท dependency test

Explain a change in source data

  1. Add is_adult to the member repeat.
  2. Enter one member age 17 and inspect the calculated value in preview or a test export.
  3. Change the age to 18; predict and observe the new value.
  4. Move the calculation outside the repeat as a thought experiment. Explain why a single ${member_age} reference is ambiguous without an aggregate or indexed reference.
Expected result

The same member changes from 0 to 1. Outside the repeat, use a repeat-aware function such as count-if() rather than assuming one member's age represents the household.

Specify the output before writing the expression

For each calculated field, write five lines in the source review: what it represents, its unit or type, which fields it reads, when it may be blank, and what changes it. For adult_count, the unit is people, the value is a nonnegative integer, and it reads all member ages in one household submission. Its value should change when an age or roster row changes. For rand_draw, the value is a number in [0,1), drawn once for a submission and intentionally stable after edits. Those two fields need different recalculation behavior even though both use calculate.

FieldDepends onCorrection test
is_adultOne member's age17โ†’18 changes 0โ†’1
adult_countAll member agesRemove an adult row; total falls by 1
publish_okConsentYesโ†’No changes 1โ†’0
rand_drawNew submission eventReopen and edit; draw stays fixed

A calculation's placement also matters. A row inside a repeat runs in that repeat's context. A row outside a conditional group may still compute when the group is hidden unless its own relevance or expression handles that case. Test the stored value after a branch change.

Write a calculation contract before the expression

Consider a field called adult_count. โ€œCount adultsโ€ is not enough to program it safely. State whether an adult is a member aged at least 18 at the listing interview, whether a missing age counts as younger than 18 or remains unknown, and whether a deleted roster row should disappear from the total. Then specify the result type as a whole number and the consumers as the follow-up rule and export. This short contract prevents a formula from quietly making a study decision.

Input stateExpected stored resultReason to test
No roster rows because consent is NoBlank in the consent-only pathNo roster was collected
One age 17 and one age 181Threshold is inclusive at 18
Age 18 corrected to 160Dependent value must recompute
Age missing in an otherwise collected rosterResolve the missing input before releaseA zero can hide missing source data

Write the expected result before previewing. Otherwise the observed result tends to become the assumed rule.

Separate a display calculation from a stored decision

A note can show a collector a calculated total, but the workflow also needs a stable field name in exports and publishing. Store the total in adult_count, then display it with a reference in a note. The field name belongs in the data dictionary. The note's wording can change without changing the data contract. If the total controls follow-up eligibility, add a second named field, such as followup_eligible, whose expression applies the full rule. That makes the count and the decision separately inspectable.

Trace the chain after each edit: age โ†’ adult flag โ†’ count โ†’ eligibility โ†’ relevance or publishing filter. When the final visible screen looks correct, inspect stored values too. A hidden field can be wrong while the collector sees no warning.

Knowledge check

Module 6 check

Attempts remaining: 3
1. Where is a calculated value normally stored?
2. Which type recalculates when dependencies change?
3. When is calculate_here appropriate?
Module 7 ยท Part II ยท Calculations and logic

Calculate with incomplete answers

Missing values, conversion, and safe defaults

By the end of this module, you can:
  • Distinguish zero from blank
  • Guard a division
  • Test unavailable inputs

Define what blank means before calculating

A count of zero means the respondent reported none. A blank may mean not asked, refused, unknown, or not yet entered. These values should not be combined without a rule. If a calculation turns blank into zero, the downstream dataset may make an incomplete interview look like a true zero. If a calculate field drives a publishing filter, that difference can change whether a row is published at all.

InputsInterpretationExpected derived result
Adults 0, total members 3Observed zero adultsAdult proportion 0
Adults blank, total members 3Adult count unavailableBlank or explicit missing code
Adults 2, total members 0Contradictory answersStop or flag; do not divide
Adults 2, total members blankDenominator unavailableBlank or explicit missing code

Guard calculations at their inputs

Use empty(${field}) to test whether a field has no value. For a rate, check that the denominator is present and greater than zero before division. SurveyCTO uses div for numeric division. A teaching expression is if(${hh_size} > 0 and not(empty(${adult_count})), ${adult_count} div ${hh_size}, ''). Compile and test it on your server version, including an unanswered numerator, a zero denominator, and an edited denominator. For a production study, decide whether a blank result is allowed or whether the form should stop earlier with a required answer or a constraint.

Do not rely on type conversion by accident. If a source field is text because it preserves an ID, keep it text. If arithmetic is intended, use a numeric input or a documented conversion and test nonnumeric text. If a calculated value will publish into a dataset, confirm the dataset column's expected type and what blank does to an existing row.

Make a test table part of the form specification

A calculation is not complete when the formula compiles. The form specification should show test inputs, expected output, and why. For adult_count, use a three-person household with ages 17, 18, and 65; the expected count is 2. Correct age 18 to 16; expected count becomes 1. Remove the third member; expected count stays 1. Then change consent to No and inspect whether the derived count is blank or zero. That last result depends on where the calculate field sits and how relevance is applied.

Practice 7 ยท missingness

Write a safe adult-proportion rule

  1. Write the intended meaning of an empty numerator and a zero denominator in words.
  2. Write an expression that returns a proportion only when both inputs can be used.
  3. Run the four rows of the table above. Record the value stored, not only the number displayed.
  4. Explain whether your downstream publishing link should replace an existing dataset value with blank.
Model answer

A zero numerator is a real 0. A blank numerator is unknown or not asked. A zero or blank denominator should not be used in division. Publishing an empty replacement into a current-state dataset needs an explicit policy.

Try the missing-value decision

The small model below represents the rule โ€œcalculate adults divided by household size only when both values are present and the size is greater than zero.โ€ Clear an input to represent missing data; enter zero to represent an observed zero. The model runs in this course file and does not submit data.

After testing, explain whether the form should allow adult_count > hh_size. A calculation guard avoids division errors; a separate constraint or supervisor check addresses contradictory answers.

Three different meanings of zero

For a ratio such as adults divided by listed members, zero can mean โ€œno adults,โ€ โ€œno members,โ€ or โ€œthe denominator was never collected.โ€ Only the first is a valid numeric share. The expression must therefore check denominator presence and value before division. Decide whether the output for the other cases is blank, a separate status code, or a validation error; do not silently return zero for all three.

Adult countMember countAdult shareStatus
040Defined proportion
240.5Defined proportion
00BlankUndefined denominator
Blank4BlankSource count unresolved
2BlankBlankRoster total unresolved

The denominator check is part of the measure's definition, not merely a way to avoid an error message. Use a separate text or select-one status if analysts need to distinguish zero denominator from missing source answers.

Check type conversion at the boundary

Survey form values can be blank, text, integer, decimal, or a choice code. A label such as โ€œYesโ€ is not the stored choice value. Before applying arithmetic, inspect the field type and the code that will actually appear in a submission. If a preload CSV stores adult_count as text, verify that the intended conversion yields a number and decide what happens when the cell is blank or contains an unexpected word. A formula that works for the clean training row may fail for the first malformed update.

In a test export, compare input columns with every calculated output. Do this for one normal case, one boundary case, one blank input, and one corrected input. Record the workbook version with the results so the test can be rerun after a change.

Knowledge check

Module 7 check

Attempts remaining: 3
1. What is the key difference between 0 and a blank count?
2. Before dividing by a reported household size, what should you test?
Module 8 ยท Part II ยท Calculations and logic

Dates, age, and time

Time-dependent expressions and eligibility

By the end of this module, you can:
  • Choose an age rule
  • Check date boundaries
  • Record one-time timestamps deliberately

Choose the definition of age first

โ€œAgeโ€ may mean reported completed years, years since a recorded birth date on the interview day, or years at a fixed eligibility date. These are different measures. A rough expression such as int((today() - ${dob}) div 365.25) can be useful for an approximate age, but near a birthday it may not match an exact completed-years rule. For a study boundary at 18, specify whether to ask age directly, verify a date of birth, or calculate against a fixed reference date. Test the day before, on, and after the birthday.

ScenarioInput to varyQuestion to answer
Direct reported age17 versus 18Does eligibility change at the intended threshold?
Date of birthBirthday tomorrow versus todayDoes the result use completed years?
Follow-up visitVisit after a birthdayShould eligibility be frozen from baseline or updated?

If eligibility must remain tied to a baseline assessment, publish or preload that baseline result. Do not let today() silently change an earlier decision when a form is reopened months later.

Distinguish live time from captured time

today() and now() are dynamic. duration() measures total time spent filling or editing a submission. A regular calculation using a dynamic value may change as the form is edited. If the task is โ€œrecord the first time the collector reached this section,โ€ use calculate_here with an outer once() around a formatted time expression. Keep that field's meaning in its name, such as first_reached_followup_at.

type             name                         calculation
calculate_here   first_reached_followup_at    once(format-date-time(now(), '%Y-%b-%e %H:%M:%S'))

Test the device clock and time zone assumptions. Record whether the value is intended as an audit signal or an eligibility input. Audit fields should not be presented as a substitute for server receipt time.

Use time fields in a testable sequence

For a callback, the form may store a scheduled date. A calculate field could flag whether it is in the past, but that flag should update if the scheduled date changes. A one-time timestamp should stay fixed. Put both cases in a test table and edit the source date after initial entry. If the rule's intended behavior differs from what you observe, change the field type or expression rather than accepting the mismatch.

Practice 8 ยท boundary dates

Specify an eligibility and timestamp test

  1. Write whether your capstone defines adult eligibility from reported age or date of birth.
  2. Test ages 17 and 18, and one correction between them.
  3. If using a date of birth, test the day before and the day of the 18th birthday against a fixed reference date.
  4. Draft a one-time checkpoint timestamp and state when it should stay unchanged.
Decision rule

Do not use an approximate date formula for a legal or study-critical threshold without validating the boundary. A one-time checkpoint value should not recalculate on return to the section or on later edit.

Do not let device time stand in for study time

For a household eligibility rule, the relevant date might be interview day, baseline date, or a fixed study cutoff. A device's today() answers only what that device currently considers today. If a collector's clock is wrong or an interview is reopened after the cutoff, the result may change. Write the reference date into the form or preload it from an approved source when the rule must be fixed. Store both the source date and the derived eligibility flag so an analyst can reproduce the decision.

Use a test with a person born exactly 18 years before the cutoff, one day before that date, and one day after. Repeat after changing the device date in a controlled training setting or by substituting a fixed reference-date input. If the result is used for case routing, check that the published flag and case list match the form's displayed decision.

Knowledge check

Module 8 check

Attempts remaining: 3
1. Why can a rough age-from-year rule be wrong near a birthday?
2. Which field type captures a value when the form reaches its location?
Module 9 ยท Part II ยท Calculations and logic

Random values and stable assignment

once(), random(), and reproducible checks

By the end of this module, you can:
  • Store a draw once
  • Separate draw from assignment
  • Test edits and reopening

Separate the draw from the assignment

A random assignment needs a stored draw and a rule that maps it to a group. Put once(random()) alone in a calculate field called rand_draw. Then use another calculate field for if(${rand_draw} < 0.5, 'A', 'B'). A draw below 0.5 maps to A; a draw from 0.5 up to but not including 1 maps to B. The threshold is a design decision, not proof that a small pilot will contain exactly half in each group.

type        name          calculation
calculate   rand_draw     once(random())
calculate   assignment    if(${rand_draw} < 0.5, 'A', 'B')

SurveyCTO warns against putting random() directly in relevance or nesting once(random()) inside another expression. A new draw could occur when a form is edited. Use the stored draw in later expressions, constraints, and publishing fields.

Test stability, not just balance

  1. Complete a synthetic interview and record the stored draw and group.
  2. Move backward and change an unrelated answer. Check the draw and group again.
  3. Save, reopen, edit, and check again.
  4. Create a different new interview. It should have its own draw; do not expect a particular group.

If the study has stratification, quotas, or central assignment, this simple threshold may not be the design. A server dataset can hold authoritative assignments, or a different randomization workflow can be used. The form must implement the approved assignment method and record enough information to audit it.

Calculation drill

Enter the calculation for the dedicated random-draw field.

Practice 9 ยท stable assignment

Write an assignment audit

  1. Add two calculated fields to a practice copy: rand_draw and assignment.
  2. Predict the group for synthetic draws 0, 0.4999, 0.5, and 0.9999.
  3. Run the edit and reopen test. Keep the draw and group in the test log.
  4. Explain why re-randomizing after an interview correction would be a problem.
Expected groups

0 and 0.4999 map to A; 0.5 and 0.9999 map to B. The saved draw must remain stable when other answers change.

Audit an assignment as a stored decision

In a training export, keep rand_draw, assignment, form version, and submission KEY together. Recalculate the expected group from the draw in a separate worksheet or script and compare it with the stored assignment. A mismatch might indicate an expression change, a threshold error, or an edit that recomputed one field but not the other. Do not judge a small sample by whether its A/B counts are exactly equal; check the rule for each record.

If a design needs a centrally controlled assignment, a form-local random draw may be inappropriate. A server dataset can hold an approved assignment by household ID; the form preloads it and stores the ID and assignment used. That changes the workflow: missing IDs and stale attachments must be tested before fieldwork. Choose the approach from the study protocol, then document it in the form specification.

Knowledge check

Module 9 check

Attempts remaining: 3
1. Where should once() appear in a random draw?
2. What should happen to an existing assignment after a form is reopened?
Module 10 ยท Part II ยท Calculations and logic

Repeats and aggregate calculations

Rosters, index(), count-if(), and sum()

By the end of this module, you can:
  • Build a member repeat
  • Calculate adult counts
  • Test a changed roster

A repeat creates many records within one submission

Use a begin repeat and end repeat pair for a roster. Each instance describes one member. A fixed repeat_count can come from household size; alternatively, the collector can add members manually. Decide which is safer for the study. If the count is later reduced, test how the device handles already-entered member rows. Keep a stable person identifier if the same person must be found across later forms; the current row position alone may change.

type            name          label                       repeat_count
integer         hh_size       Usual residents
begin repeat    members       Household member            ${hh_size}
text            member_name   Member name
integer         member_age    Age in completed years
calculate       member_order                              index()
end repeat      members

index() gives the current repeat position. It is useful for a roster slot, but not a durable person ID if rows can be reordered or removed.

Aggregate across all instances

Outside the repeat, count(${members}) returns the number of repeat instances. count-if(${members}, ${member_age} >= 18) counts adults. sum(${member_income}) totals a numeric field in the repeat. SurveyCTO also has conditional aggregate functions such as sum-if(). Use the function that matches the intended unit. A single ${member_age} reference outside a repeat is not a household summary.

Memberscount()count-if(age >= 18)
17, 18, 6532
17, 16, 65 after correction31
17, 16 after removing one row20

For a later form, joining by HH003 plus a person ID is safer than joining only by household ID and roster position. A deletion can make the second row refer to a different person.

Test the edit path and the export shape

Create three roster rows, correct an age, and remove one row. Check that the aggregate follows each change. Then submit a synthetic record and inspect the exported repeated data. Keep the household identifier in the repeated-data context needed for analysis. If a later publishing link sends member data to a dataset, decide whether one wide household row or one long row per person is needed; Module 18 develops that choice.

Aggregate drill

Write a calculation for the number of members aged 18 or older.

Practice 10 ยท roster arithmetic

Build and correct a three-person roster

  1. Enter ages 17, 18, and 65. Predict the adult count before checking it.
  2. Change 18 to 16. Confirm the count updates.
  3. Remove the 65-year-old row. Confirm the count updates again.
  4. Describe how you would identify the same member in a follow-up form after the roster order changes.
Expected results

The adult counts are 2, then 1, then 0. Use a stable person ID for longitudinal linking; index() alone is a position.

Change a roster and watch the total

This model mirrors the adult-count test. Clear an age to represent an unanswered member. The count includes entered ages of at least 18; the form's real required rule should prevent an unanswered age from being finalized when that is the study policy. Try 17, 18, 65, then change 18 to 16.

A displayed count here is a learning model. The workbook's count-if() expression and the server preview remain the authoritative test of the actual form.

Knowledge check

Module 10 check

Attempts remaining: 3
1. Which function counts adult members in a repeat?
2. Which function identifies the current repeat position?
Module 11 ยท Part III ยท Field practice

Collect and update forms

Web, Collect, offline state, and new versions

By the end of this module, you can:
  • Run an offline cycle
  • Locate unsent records
  • Plan device updates
Build on Module 1: You already sent one practice interview and checked its CSV export. Here you deliberately disconnect a real test device and prove what survives locally, what must be sent later, and how a revised deployed form reaches collectors.

Know the state of a field record

A form can be completed in a web browser or in SurveyCTO Collect. In Collect, an interview can be open, saved but unfinished, finalized, ready to send, sent, and eventually visible on the server. The words โ€œcompletedโ€ and โ€œuploadedโ€ are not interchangeable. For a missing household, check the device state before deciding the server lost it. For a web form, distinguish a preview run from a deployed form submission.

ObservationWhat it provesWhat it does not prove
Preview passedTested draft paths workedDeployment and device delivery
Finalized locallyCollector finished the local formServer received it
Sent on deviceDevice attempted and recorded sendingCorrect downstream publishing
Submission visible on serverServer received a recordDataset row is current

Test an offline cycle on the actual device

  1. While connected, configure the practice server and download the deployed form.
  2. Disconnect. Complete a synthetic household interview, then finalize it.
  3. Confirm it remains stored on the device and note its local status.
  4. Reconnect, send, and verify the server received it exactly once.
  5. If the form is attached to a dataset, refresh the relevant attachment or case list and verify the new row appears where expected.

Do not clear device storage or remove a form to โ€œfixโ€ a send problem while local records remain. Record the device, user, form version, affected IDs, and error. A supervisor can then troubleshoot without losing unsent work.

Distribute changes as a field procedure

A deployed new version is not automatically the version on every offline device at that moment. A release instruction should say when to connect, where to obtain updates, whether saved interviews can be safely reopened, and what test confirms the device has the intended version. If the form uses a dataset attachment, updating the form and refreshing the data source may be separate tasks. A collector should never infer that a new case is absent from the project merely because the local case list has not refreshed.

Practice 11 ยท device-state log

Trace one offline record

  1. Use an invented ID HH004. Record the form version before disconnecting.
  2. Complete and finalize while offline. Write down the device state.
  3. Reconnect, send, and find the server submission. Record its KEY.
  4. Describe how you would distinguish a slow send from a duplicate if the collector considered entering HH004 again.
Expected evidence

The device state, server submission list, and any downstream dataset row are separate evidence. Match records by ID and timestamp before re-entry.

Knowledge check

Module 11 check

Attempts remaining: 3
1. A finalized record in Collect is absent from the server. What should you check?
2. A saved form is updated on the server. What should the team test?
Module 12 ยท Part IV ยท Connected data

Attach and preload data

CSV versus server dataset inputs

By the end of this module, you can:
  • Identify a lookup key
  • Use pulldata()
  • Test missing and stale records

Choose a source that matches the update need

A direct CSV attachment is useful for a fixed sample or a small training file. A server dataset is useful when the lookup needs to be updated from multiple forms or maintained centrally. Both become local preloaded data when attached to a form. The form needs the correct columns, a stable lookup key, and a refresh procedure for field devices. An attachment is an input to a form; publishing is a separate rule that writes data into a dataset.

SourceGood training useUpdate test
households_csv.csvThree known household IDs for lookup practiceReplace attachment and obtain new form/support file
hh_registry_training datasetListing submissions update current household rowsPublish a new row, then refresh the attached copy

Read every argument of pulldata()

pulldata('hh_registry_training', 'head_name', 'hhid_key', ${hhid})

The first argument is the attached dataset ID (or CSV basename); the second is the column to return; the third is the column used to find a row; and the fourth is the form's lookup value. If HH003 is in the form but the dataset row uses hhid_key=HH3, no match is found. If two source rows share the same supposed unique key, the lookup is ambiguous; fix the source rather than guessing which row the form will use.

Preloaded values come into expressions as text strings even when a source column looks numeric. Convert deliberately before arithmetic. Keep the raw pulled value and a converted calculated value separate when troubleshooting a mismatch.

Test missing and stale lookups

  1. Attach the training CSV or dataset to the follow-up form.
  2. Use HH001 and confirm the expected head name appears.
  3. Use HH999 and confirm the form follows an explicit missing-ID path.
  4. Change the source name for HH001, then check the old device before refreshing and after refreshing.

A server dataset can be current while a device still has an older attached copy. Ordinary server publishing requires server receipt, and refreshed dataset files may not be immediately available after every update. Record the actual delay observed in your training test and use a field procedure that accommodates it.

Lookup drill

Write the expression to pull head_name using the entered hhid.

Practice 12 ยท preload

Find a match and a miss

  1. Inspect the included household CSV and identify its unique key.
  2. Write the pulldata() expression for the head name and explain each argument.
  3. Test HH001 and HH999. Record the pulled value and the form's response to a miss.
  4. Write one sentence for a collector explaining what to do when the selected case is missing from the attached data.
Model answer

The source has one row per hhid_key. A missing result is a data or refresh issue to investigate; the collector should not invent a head name.

Knowledge check

Module 12 check

Attempts remaining: 3
1. What does pulldata() need to find a row?
2. If the server dataset changed but a device is offline, what might the device use?
Module 13 ยท Part IV ยท Connected data

Build dynamic choice lists

search() and filtered options

By the end of this module, you can:
  • Choose stable choice values
  • Filter a long list
  • Test zero and multiple matches

Use a dynamic list when fixed choices would be unwieldy

A long school, village, or household list can be loaded from an attached CSV or dataset. In SurveyCTO, a search() expression in the appearance column selects source rows. The matching choice row on choices specifies which source column supplies the stored value and which supplies the displayed label. The stored value should be a stable ID, never merely a name that can be duplicated.

survey:  type=select_one households
         name=selected_hhid
         appearance=search('hh_registry_training', 'matches', 'village_code', ${village_code})
choices: list_name=households  value=hhid_key  label=head_name

This example lists households in the selected village. The value and label cells on that special choices row name source columns; they are not literal household values. Verify the exact source headers and the attached dataset ID before uploading.

Choose an appropriate match and filter

search(source) can list all distinct rows. matches tests an exact value in a named column; contains, startswith, and endswith support other searches. You can also filter by another column. A broad text match may show several people with the same name, so display enough context, such as household ID and village, for a collector to choose correctly. A search that returns zero rows is an expected test case, not an excuse to fall back to the first row.

InputExpected listCheck
Village V01Only V01 householdsIDs and labels are paired correctly
Village V99No rowsCollector gets a clear missing-list procedure
Two households named AnaTwo distinct IDsCollector can distinguish them

Keep list updates controlled

A list can change when the dataset refreshes. If a code used in an existing submission disappears or changes meaning, analysis and edits become harder. Define who owns the source, when it refreshes, which column is the stable key, and how a change is communicated to collectors. On an offline device, confirm the new list actually arrived before asking a collector to find a newly added household.

Practice 13 ยท dynamic choices

Design a village-filtered selector

  1. Choose a stored value and a display label from the included household lookup file.
  2. Write a search() appearance that returns only rows in the selected village.
  3. Test V01, V99, and two identical display names with different IDs.
  4. Explain why the published follow-up should store the household ID even if the UI shows a name.
Expected design

Use hhid_key as the selected value, and a label that helps distinguish rows. Treat no match as a controlled exception.

Knowledge check

Module 13 check

Attempts remaining: 3
1. What should be stable in a dynamic choice list?
2. A search returns no matches. What should the form do?
Module 14 ยท Part IV ยท Connected data

Design a server dataset

Rows, unique IDs, ownership, and refresh

By the end of this module, you can:
  • Define a dataset key
  • Choose data columns
  • Separate source submissions from current state

Model the dataset before configuring links

A server dataset is a managed collection of records on the SurveyCTO server; its contents can be uploaded or downloaded as a CSV. For the running example, hh_registry_training stores one row per household. Its key is hhid_key; other columns might be head_name, village_code, adult_count, last_outcome, and last_visit_date. Write the meaning and owner of each column. Decide whether the dataset represents a historical event log or the latest state of each household. This example is a latest-state lookup, so repeated submissions should update the same household row using a stable key.

ColumnMeaningSource
hhid_keyUnique household IDListing and follow-up forms
head_nameName shown in follow-upListing form
adult_countCurrent listed adult countListing calculation
last_outcomeLatest follow-up statusFollow-up form

Use a unique ID field when rows must update

SurveyCTO can enforce that a nominated Unique ID field contains unique, nonblank values. When you configure a form to publish to that dataset, map the form's identifying field into the dataset's unique column and choose it as the field to identify unique records. If you omit the join, each submission can append a new row instead of updating the existing household. If a form publishes an empty or misspelled ID, stop and investigate before trying to merge records by name.

The raw listing and follow-up submissions remain separate evidence. The dataset is a maintained view for preloading or workflow state. If an incorrect publish overwrites a current-state column, use submission history and a correction plan to reconstruct the intended value.

Plan refresh and ownership

Dataset contents can be edited or uploaded manually, published from forms, or updated programmatically. For manual uploads, distinguish append, merge, and replace. A replace operation on the wrong file can discard current rows; use the synthetic training dataset and keep a copy before trying it. Dataset downloads and attached copies may lag recent server updates, so include refresh time in a handoff test rather than assuming instant availability.

Practice 14 ยท data model

Specify the household registry

  1. Write the dataset ID, unique column, and five columns you need for the follow-up.
  2. Mark which form owns each column.
  3. Predict the row count after listing HH001 and HH002, then listing HH001 again with a corrected adult count, assuming a matching join is configured.
  4. State what should happen if a submission has no household ID.
Expected state

The dataset should have two rows, one for each ID, if the corrected HH001 submission upserts that row. A blank ID is not a valid update key and needs investigation.

Separate a source of truth from a working view

Raw listing submissions are events. The registry row is the current value the follow-up form reads. That distinction matters when a correction arrives. If two records for HH001 disagree on head name, the registry needs a documented rule for which value wins. โ€œLatest arrival winsโ€ can be wrong when an older offline interview is sent late. For a high-stakes workflow, include a reviewed update process or explicit event time and conflict procedure.

Keep the registry narrow. Publish only the fields the next form needs, such as household ID, display name, village code, and a non-sensitive routing flag. An analyst can still use the full submissions for research. A field device should not receive a broad dataset merely because it was convenient to map every form field.

Knowledge check

Module 14 check

Attempts remaining: 3
1. A server dataset stores which structure?
2. What does a unique ID field support?
Module 15 ยท Part IV ยท Connected data

Publish a form into a dataset

Mapping, action, filter, and timing

By the end of this module, you can:
  • Configure Publish into
  • Map and verify fields
  • Trace a received submission

Configure the link from the destination dataset

On the server console's Design tab, find or create the destination dataset. Choose Publish into on that dataset and select the source form. Map each form field to the intended dataset column. Do this field by field for the training example instead of choosing Add all; the explicit mapping makes the data contract visible. Only a field that was relevant in the submission can publish its value.

Listing form fieldDataset columnReason
hhidhhid_keyStable join key
head_namehead_nameFollow-up display
village_codevillage_codeField routing and search
adult_countadult_countCalculated summary

Choose the link's other settings deliberately

Choose wide format for one dataset row per household submission. Select hhid as the form field that identifies unique records so a later listing for HH001 updates the HH001 row. Add a calculated filter field when only eligible or completed records should publish; the filter must equal 1. Decide whether to select Publish existing data. If selected, earlier submissions can be backfilled into the new link. If not, the link starts with future submissions. Write down which choice you made.

For encrypted forms, only fields configured as publishable can flow into a dataset. A training server with invented data simplifies the first exercise, but a real project must decide which fields may be shared before configuring the link.

Test the link end to end

  1. Submit HH001 with consent Yes, two adults, and a head name.
  2. Wait for server receipt and locate the submission KEY.
  3. Inspect the publishing link status and dataset row for HH001.
  4. Submit a second HH001 record with a corrected adult count.
  5. Confirm whether the same dataset row updates and whether its value matches the chosen update action.

When a dataset row is absent, check in order: the source submission was received; the publishing link is enabled; the filter equals 1; the mapped source fields were relevant; the join field is nonblank; and the dataset view has refreshed. Do not change several settings at once while diagnosing.

Practice 15 ยท publishing map

Write a four-column contract

  1. Copy the field mapping above into a release note.
  2. State the join field, format, filter policy, and Publish existing data choice.
  3. Run HH001 and an HH001 correction. Record both submission KEYs and the resulting dataset row.
  4. Explain why two source submissions can correspond to one current-state dataset row.
Expected result

The submissions remain two historical records. With a join on household ID, the dataset can show one updated row for HH001.

Use a publishing-link test record

Give each new link a known synthetic case. For HH004, write the expected dataset row before sending the form: hhid_key=HH004, head name โ€œTest Personโ€, village V02, adult count 1. After the server receives the submission, compare every mapped column with this expectation. If the form uses review and correction, note whether the record was held and approved; an awaiting-review record may be received but not yet published.

Keep a screenshot or written export of the link's map and settings. The workbook alone cannot show which destination dataset, update action, filter, or Publish existing data choice was configured in the console. That configuration is part of the maintained source for a connected workflow.

Walk one record through the publishing link

Suppose HH001's submitted listing has head_name=Ana, adult_count=2, and publish_ok=1. The publishing link reads the received submission, applies its filter, maps hhid to hhid_key, and writes the mapped values to the destination dataset. A second HH001 listing with corrected adult_count=3 leaves both submissions available as source records but should update the one current HH001 row when the join and Replace action are configured as specified. The dataset is therefore an operational view with a documented rule, not a substitute for submission history.

CheckpointEvidence to recordFailure it isolates
Server receiptForm ID, version, submission KEY, receipt timeUnsent or wrong-form record
Filterpublish_ok value in that submissionCorrectly excluded record
MappingSource field and destination column pairsBlank or wrong column
JoinHH001 key and destination unique-ID settingDuplicate or failed update
DestinationRow value after refreshStale view or failed link

Keep this trace beside the field map. It is more useful than a screenshot of a successful Publish button because it identifies which record and settings produced the row.

Decide how existing data enters a new link

When a publishing link is created after submissions already exist, record whether Publish existing data was selected. If you backfill, inspect the resulting row count and a few named keys before allowing downstream forms to use the dataset. If you start with future submissions only, an earlier HH001 submission will not appear solely because the link now exists. A lookup test using that old record could therefore fail even though the new link works.

For the capstone, make one submission before creating the link and one after. Predict the registry under each setting, then verify. This deliberately tests the backfill decision and avoids confusing it with an attachment or pulldata() problem.

Knowledge check

Module 15 check

Attempts remaining: 3
1. Which console action creates a form-to-dataset publishing link?
2. What can Publish existing data do?
Module 16 ยท Part IV ยท Connected data

Keys and update actions

Append, replace, add, concatenate, and filters

By the end of this module, you can:
  • Choose an upsert key
  • Predict two submissions
  • Test duplicate and absent keys

Predict a dataset before clicking Publish

A publishing link can append a new row for every included submission or use a form field as a join key to update an existing row. For a household registry, hhid is the join. Map it into hhid_key. If HH001 is published twice, the intended current-state dataset has one HH001 row. Without a join, it may have two. A head name is an unsafe join because two households can share a name and a spelling correction can change it.

Received submissionJoin fieldDataset action
HH001, adult count 2HH001Insert HH001
HH002, adult count 1HH002Insert HH002
HH001, adult count 3HH001Update HH001, if the link uses this key

Choose a field-level update action

The default Replace action uses the new mapped value as the dataset value. Add increases an existing numeric value by the submitted value. Concatenate appends or prepends text with a separator; the published text has a length limit, so it is not a safe unlimited event log. For adult_count, Replace is usually appropriate because a corrected count is the new total, not an increment. For a cumulative tally, Add may be appropriate if repeated submissions are guaranteed not to be replayed. Do not use Add for a value that can be resent or edited unless duplicate protection is designed.

Predict HH001 after a correction

Existing adult_count=2; new submission reports 3. Choose the action to see its effect.

Choose an action.

A publishing filter is a calculated decision

A publishing link can include a form submission only when a chosen field equals 1. A calculate field such as publish_ok = if(${consent} = 'yes', 1, 0) makes the rule explicit. Decide whether refusal records should create a lookup row, update a case state, or remain only in raw submissions. Those are project decisions; no single filter fits all studies. Test 1, 0, and blank. Also decide whether a changed answer and corrected submission should remove or revise a row published earlier; a filter on later submissions does not automatically describe the handling of old dataset state.

Practice 16 ยท update prediction

Make a before-and-after table

  1. Start with HH001 adult count 2 and HH002 adult count 1.
  2. Predict the dataset after HH001 reports 3 under Replace, Add, and no join.
  3. Write a 1/0 publish filter for consent Yes and test Yes, No, and blank.
  4. Explain how you would detect a duplicated submission before using Add for any measure.
Answer key

Replace gives HH001=3; Add gives HH001=5; no join can add another HH001 row. A filter such as if(${consent} = 'yes', 1, 0) includes Yes and excludes No or blank under the stated rule.

Simulate an update before configuring a link

Start with a dataset row for HH001 whose adult count is 2. Enter a new submitted count, select an update action, and decide whether the publishing filter passes. A no-join link appends a second row; Replace and Add use the matching household key. This model intentionally shows why adding a corrected total is wrong.

After using the model, write the actual field map you would put in the console. The simulator cannot see your server's links, encrypted fields, or pending submissions.

Knowledge check

Module 16 check

Attempts remaining: 3
1. Without a joining field, what happens to new submissions?
2. Which filter value allows a submission to publish when a filter field is configured?
Module 17 ยท Part IV ยท Connected data

Link multiple forms

A listing form feeds a follow-up form

By the end of this module, you can:
  • Draw a two-form data path
  • Publish then attach
  • Test refresh and missing lookup

The connected two-form path

hh_listing_training
HH003, Cora, 3 adults
Publish into โ†’
hh_registry_training
hhid_key=HH003
Attach to โ†’
hh_followup_training
pulldata by HH003

The listing form does not directly send a value to the follow-up form. It sends a submission to the server. A publishing link maps that submission into a dataset row. The follow-up form has the dataset attached and pulls the intended row by the stable ID. A separate cases dataset can expose the follow-up form to an assigned collector. Make each handoff explicit; otherwise a missing name in follow-up can be misdiagnosed as a calculation error when the dataset simply did not refresh.

Configure in an order that produces evidence

  1. Upload and deploy the listing form. Verify its form ID and synthetic test submission.
  2. Create hh_registry_training with hhid_key as the unique ID field and the columns the follow-up will need.
  3. Use Publish into to map listing fields. Send HH003 and inspect the dataset row.
  4. Upload the follow-up form and attach hh_registry_training as its preloaded source before deploying and testing it.
  5. In the follow-up, use the case ID or a selected household ID to pull head_name. Test HH003 and a missing ID.
  6. After publishing a correction from the listing form, refresh the attached dataset on the intended device and check the new value.

This order helps distinguish a form error, publishing error, attachment error, stale local copy, and lookup-key mismatch.

Write a connection contract

Contract itemTraining valueFailure test
Source formhh_listing_trainingWrong form chosen in Publish into
Datasethh_registry_trainingWrong dataset ID attached
Joinhhid โ†’ hhid_keyHH003 versus HH03
Lookuppulldata(..., 'hhid_key', ${hhid})HH999 missing
RefreshCollector obtains current attachmentOld name still shown

Store this contract next to the maintained XLSForms. The console settings are part of the field system even though they are not in either workbook.

Practice 17 ยท linked-form handoff

Trace HH003 through every layer

  1. Draw the path from listing answer to follow-up display, naming the two forms and the dataset.
  2. For each arrow, write the exact evidence you would collect in a live test.
  3. Change HH003's head name in a new listing submission. Predict what the follow-up shows before and after an attachment refresh.
  4. Write a missing-record procedure for HH999.
Expected sequence

Verify source submission, publishing link, HH003 dataset row, attached copy, and follow-up pulldata() result. A stale attachment can continue to show the earlier value until refreshed.

Trace a failed connection without guessing

First known good pointNext artifact to inspect
Listing form preview passesDeployed version and received synthetic submission
Submission receivedReview status and publishing link
Link ranHH003 row and mapped columns in registry
Registry row correctFollow-up attachment and its refresh time
Attachment currentpulldata() source ID, lookup column, and HH003 value

Stop at the first mismatch. If the registry row is absent, changing the follow-up form will not create it. If the row is correct but the device shows an old value, repeating the listing interview can create duplicate raw records without fixing the attachment. The table turns โ€œthe forms are not connectedโ€ into a specific, reproducible fault.

Keep three copies of a household record straight

The two-form workflow can show three different states at one moment: the latest listing submission on the server, the current registry dataset row, and the dataset copy attached to a collector's form. They are related by publishing and refresh steps but are not one object. If the head name was corrected to โ€œAna Mariaโ€ in a new listing, first check that the submission reached the server. Then check the registry row. If it is current there but the follow-up still shows โ€œAna,โ€ update the attached data on the collection device and retry the lookup.

Latest submissionRegistry rowDevice lookupLikely next check
OldOldOldSend or locate corrected submission
NewOldOldPublishing filter, map, join, link status
NewNewOldRefresh attached dataset on device
NewNewNewRecord successful end-to-end test

Do not edit the follow-up expression to compensate for a stale attachment. Fix the state that is stale.

Specify the missing-key path

A collector may mistype HH001 as HH01, select a case whose registry row has not published, or work with an outdated case list. A blank pulldata() result cannot tell those causes apart. The form should expose the entered or selected key, show a clear โ€œhousehold not foundโ€ instruction, and prevent an ordinary completed follow-up until the ID is resolved. The field procedure should say whether the collector refreshes attached data, calls a supervisor, or records a separate unresolved-contact event. Never fill a household's name by guessing from a near match.

Test both HH001 and HH999. For HH999, verify the displayed message, the fields that remain accessible, and what a submitted unresolved record would contain. That negative path is part of the linked system, not an optional edge case.

Knowledge check

Module 17 check

Attempts remaining: 3
1. What is the order for making listing data available in a follow-up form?
2. What should both forms use to identify the same household?
Module 18 ยท Part IV ยท Connected data

Repeated publishing and offline updates

Wide, long, and device-local transitions

By the end of this module, you can:
  • Choose wide or long
  • Specify a repeated key
  • Distinguish server and offline publishing

Choose wide or long from the destination unit

In wide publishing, one form submission produces one dataset row. Repeat fields become numbered columns such as member_age_1, member_age_2, and so on. This can work for a household summary with a limited roster. In long publishing, each repeat instance produces its own dataset row. This is often better when another form needs to search individual members. Long publishing requires a field inside the repeat that uniquely identifies each destination row; a household ID alone is not enough for several members in one household.

Destination questionFormatExample key
What is HH003's current adult count?Wide household rowHH003
What are the ages of each HH003 member?Long person rowsHH003-P01, HH003-P02

For long format, mapped repeated fields must come from the same repeat group or a suitable parent repeat. A repeated unique field must be mapped and selected for identifying records. Test adding, editing, and removing a repeat instance to see how the dataset changes.

Server publishing and offline publishing are separate plans

Ordinary server publishing follows server receipt. A device offline in the field cannot assume a new listing submission is already available as preload in its next form. SurveyCTO also supports offline dataset publishing in Collect, but that feature needs its own configuration, permissions, compatible forms, and device tests. The course's basic capstone uses server receipt and a deliberate attachment refresh; it does not silently imply instant offline updates.

If an advanced study needs two forms on one device with no network between them, write an offline test covering a new case, a corrected value, a second device, and later synchronization. When devices disagree, define which dataset state wins and how supervisors discover conflicts. A same-device demonstration alone does not prove a multi-device workflow.

Reconcile the source and destination

When a repeat publishes long, count the expected repeated rows from the form and compare them with dataset rows for that submission. Check the repeated key, household key, and selected filter. A difference may be intended if only eligible members publish, but the rule must be documented. Avoid changing a roster's row order without testing whether member identities remain stable.

Practice 18 ยท repeated publishing

Choose a destination shape

  1. For three members in HH003, sketch the wide row and the three long rows.
  2. Choose a unique key for each long row that remains stable after a roster edit.
  3. Explain why HH003 alone cannot uniquely identify all three long rows.
  4. Write separate tests for ordinary server publishing and any planned offline publishing.
Expected interpretation

Wide is one household row; long is one row per member. A long-row key must identify a member instance, not just the parent household.

Knowledge check

Module 18 check

Attempts remaining: 3
1. What does long-format publishing do with repeated records?
2. Does ordinary server publishing make new data instantly available on an offline device?
Module 19 ยท Part V ยท Case management

Cases as a fieldwork queue

Case datasets and the Manage Cases interface

By the end of this module, you can:
  • Prepare a case CSV
  • Connect a case to a form
  • Use caseid in form data

A case is the subject of work, not a submitted visit

Case management reorganizes collection around households, schools, clinics, or another tracked unit. A cases dataset supplies the list shown through Manage Cases. One case can offer one or more forms, and one case can have several submissions over time. The case ID identifies the subject; a submission KEY identifies one completed visit. Do not use a changing display label as the case ID.

Cases dataset row
id=HH003
shows โ†’
Manage Cases
label and forms
opens โ†’
Follow-up form
caseid=HH003
produces โ†’
Visit submission
new KEY

The attached workshop photo shows the original case-routing discussion. The typed diagram and tables here are the operational version to use for the exercises.

Workshop whiteboard sketch of a cases list, SurveyCTO server, user roles, and tablet case queues
Source workshop whiteboard. Use the typed case model in this module for exact field names and steps.

Prepare the cases CSV

The required columns are id, label, and formids. The ID must be unique and nonblank. The label is shown to the collector. formids lists the form IDs offered for that case, separated by commas when there is more than one. Optional columns include users, roles, enumerators, and sortby. Include the standard columns in a maintained training CSV so assignments can be added without restructuring it later.

id,label,formids,users,roles,sortby,enumerators
HH001,Household HH001,hh_followup_training,,,1,
HH002,Household HH002,hh_followup_training,,,2,

Blank user and role values make the initial training cases broadly visible under the platform's case-list rules. Before actual field assignment, replace them with approved test usernames or role IDs and verify the effective view with those accounts. Use only invented IDs and labels in the trial.

Connect caseid to the form

Include the caseid metadata field in the follow-up form. When the user opens a form from Manage Cases, SurveyCTO populates it with the selected case's ID. If the same form is opened from Fill Blank Form, caseid can be blank. The form should either require a valid case path for case-managed work or define a separate manual-ID procedure. For the training follow-up, the lookup uses the selected case ID as the household key.

On a mobile device, Manage Cases may need to be enabled in Admin Settings. Users refresh the case list when connected to obtain current cases. A missing case can mean a stale list, a visibility rule, a wrong role's cases dataset, or a form ID in formids that is not on the device.

Practice 19 ยท case list

Explain two case rows

  1. Download and inspect the included cases CSV.
  2. Add HH003 with formids=hh_followup_training and a unique ID.
  3. Predict what appears in Manage Cases for a user who can see all three rows.
  4. Open a case-managed form in your training server if available. Confirm that caseid equals the selected row's id.
  5. Open the form outside Manage Cases and compare the caseid behavior.
Expected distinction

Manage Cases supplies the selected case ID. A generic blank-form path has no selected case unless the workflow supplies an ID another way.

Choose the case list display for the work

SurveyCTO supports a tree view and a table view for cases. A tree can expose each case and its available forms; a table can show selected case columns for sorting and filtering. Use a short label that helps a collector find the right subject without exposing unnecessary private data. For a household training list, โ€œHH003 ยท V02โ€ is enough. A real study may require a more discreet label than a respondent's full name.

A case row can list several form IDs. That does not force every form to be completed or define their sequence; the operational procedure must say which form is due and when. A calculated or published status column can help, but a case list is only as current as its dataset and the device's refresh. Test an added case and a closed case on both the server and field device.

Design case IDs for the whole project

A case ID should survive a change of head name, village, assigned collector, and visit status. Use a text identifier even if the first set of IDs looks numeric; text preserves leading zeros and allows a prefix. Before loading the cases CSV, check for blank IDs, duplicates, leading or trailing spaces, and different spellings of the same ID. The same normalized value must be used in the registry lookup key and in caseid when the case opens a form.

CandidateIssueCorrection
HH001Stable codeKeep
HH001 Trailing space can break equalityTrim before loading
001 stored as a numberLeading zeros may be lostStore as text
Ana householdName can change or repeatUse a stable code; put name in label

Keep the ID cleaning rule with the cases file. A case list that looks right to a person may still fail an exact lookup.

Knowledge check

Module 19 check

Attempts remaining: 3
1. Which three columns are required for cases?
2. When is caseid automatically populated?

Official reading

Module 20 ยท Part V ยท Case management

Users, roles, and case visibility

Assign work and verify effective access

By the end of this module, you can:
  • Set users and roles intentionally
  • Test two collectors
  • Diagnose a missing case

Write the access plan before creating accounts

SurveyCTO users have roles that control console and collection permissions. A cases dataset can also filter rows by the users, roles, or enumerators columns. These mechanisms answer related but different questions: may this person use the system, may they fill this form, and which cases should appear in their queue? Do not assume that seeing a form implies seeing every case or that an administrator's case list matches a collector's.

Training roleNeeded taskTest
Collector AComplete HH001 follow-upHH001 visible; HH002 withheld if assigned elsewhere
Collector BComplete HH002 follow-upHH002 visible; HH001 withheld if assigned elsewhere
SupervisorReview submissions and case progressCan see the required records without unnecessary design rights
Form managerUpdate forms, datasets, and linksCan manage configuration and see release evidence

Use test accounts or the server's training environment. If you create named users for a class, tell them what records they can access. Do not place passwords in the cases CSV or XLSForm.

Filter a queue with the cases dataset

For a simple two-collector test, put the actual SurveyCTO username for Collector A in HH001's users cell and Collector B's username in HH002's. Keep formids the same if both should fill the follow-up form. Reload the case list as each collector. A comma-separated users cell can list more than one approved account. The roles column can list role IDs instead of named users for role-based queues. If enumerator IDs drive visibility, the cases dataset must be linked to the proper enumerator dataset and tested with the enumerator selection flow.

Do not infer the combining behavior of several nonblank visibility columns from a single admin screenshot. Keep the first lab to one filter mechanism, then test every intended combination with separate accounts before a pilot.

Diagnose a missing case in order

  1. Confirm the case row exists in the correct cases dataset and has a nonblank unique id.
  2. Confirm the current user's role points to that cases dataset; multi-team or custom roles can use a non-default dataset ID.
  3. Check users, roles, and enumerators values against the actual account and configuration.
  4. Refresh Manage Cases on the device, then confirm the current form ID appears in formids and the form is downloaded.
  5. Check whether the case appears under another test account. Record which filter changes the result.

This sequence protects against a common mistake: adding duplicate case rows because the first row is invisible under the current account.

Practice 20 ยท two queues

Test effective access

  1. Prepare a two-row cases CSV for HH001 and HH002 using test usernames, not personal passwords.
  2. Update only the users values in that synthetic CSV.
  3. Sign in as each test collector or have a colleague verify the view. Record visible and hidden cases.
  4. If a case is missing, use the five-step diagnosis rather than making another row.
Completion evidence

Keep the CSV, the role/username plan, and a result from each collector view. One admin view is not evidence that field access is correct.

Practice the simplest case filter

This model shows one mechanism at a time: the users column. HH001 is assigned to collector_a; HH002 is assigned to collector_b; HH003 has a blank user value and is visible to both in this training model. Real role and enumerator filtering should be tested separately on the server.

The model predicts which rows to expect after refresh. It does not grant permission or create users. A server role can still prevent collection even when a case is listed.

Knowledge check

Module 20 check

Attempts remaining: 3
1. What can the users column control?
2. How should a role plan be verified?
Module 21 ยท Part V ยท Case management

Case state transitions

Create, assign, follow up, and close

By the end of this module, you can:
  • Map a case lifecycle
  • Use publishing to update state
  • Test repeated and out-of-order submissions

Define a case lifecycle as data changes

A case list can be static, but a connected workflow can update it from submitted forms. Write the states first: not_started, assigned, callback, complete, and perhaps closed. Define who may move a case between them and what evidence supports each transition. A follow-up submission with outcome callback might update a case status and next-visit date; a completion might close the case or remove the follow-up form from its formids. These are publishing decisions, not automatic effects of filling a form.

Current stateNew eventIntended stateFailure to test
assignedVisit reports callbackcallback, date scheduledCallback date missing
callbackVisit completedcompleteOld callback remains visible
completeLate callback submission arrivesReview requiredComplete is overwritten by stale state

Use publishing to update a case row

A cases dataset is a server dataset with special columns. A form can publish into it. Map the stable case or household ID to its id column and use that ID as the record join. Map only the state and routing fields the form is authorized to change. If an assignment changes, update the users or roles value deliberately and test that the old and new account views reflect the handoff after refresh. A status word should not accidentally be mapped into users; the values in that column must be actual usernames if it is used for visibility.

A new case can also be generated by a publishing link, but its id, label, and formids must be supplied. Test case creation separately from case update. Keep a known-good CSV export of the synthetic cases dataset before experimenting with publishing actions.

Test order and duplicates

Submission order can differ from visit order when devices work offline. If a callback was recorded on Tuesday but sent after a Wednesday completion, a simple Replace mapping could put the case back in callback state. Decide whether the workflow should block the stale update, flag it for supervisor review, or use an authoritative timestamp and approval process. Also test duplicate sends and corrections. Dataset state is a current view; the source submissions and their KEYs remain the audit trail.

Practice 21 ยท transition table

Design a callback handoff

  1. Write the fields a follow-up form needs to publish a callback state into the cases dataset.
  2. Specify the case-row join and one assigned test username.
  3. Predict the case list before and after refresh for the old and new collector.
  4. Run a late-callback scenario after a completed visit. Write what the supervisor should see and decide.
Model answer

Join on the stable case ID. Map only approved state, date, and routing values. A late event should be detected and reviewed rather than silently reversing a completed case.

Map a handoff as a transaction

Suppose Collector A completes an initial visit and the case must move to Collector B. The transition has four observable steps: the new visit submission is received; its publish rule updates the case row's assigned user; the old account refreshes and no longer sees the case; the new account refreshes and does see it. Record all four. A success message in the form is only the first step.

Use a unique case ID for the update and publish the actual SurveyCTO username in users. Do not put a status label such as callback into the username column. If a case must remain visible to both people for a review period, list both approved usernames according to the case-management format and test both views. A supervisor should own the policy for what happens when the new collector has not yet refreshed.

Model events separately from current state

A follow-up visit is an event: HH001 was visited on a date, with an outcome and a submission KEY. A case row is current state: HH001 is assigned, callback due, completed, or closed now. Several visit events may contribute to one current row. If every follow-up submission replaces case status, a late offline callback may reopen a case after a newer completion. Write a transition rule that considers the existing state and the order of events, then test the late arrival. A simple Replace action has no inherent knowledge that โ€œcompleteโ€ should dominate โ€œcallback.โ€

Arrival orderEventNaive Replace resultDecision to specify
1Callback scheduledCallbackExpected interim state
2Interview completeCompleteClose or retain case?
3, delayedOlder callback submissionCallbackReject reversal or review manually

For a small project, a supervisor review before closing cases may be enough. For automated state changes, document the exact ordering and conflict rule and verify it on delayed offline data before using it in fieldwork.

Knowledge check

Module 21 check

Attempts remaining: 3
1. Which field should join a follow-up update to its existing case?
2. A late submission arrives after a closure update. What should you test?
Module 22 ยท Part VI ยท Operations

Monitor, review, and export

Submission QA and dataset reconciliation

By the end of this module, you can:
  • Compare submissions with published rows
  • Write an actionable query
  • Keep a reproducible export

Keep submissions and dataset rows in separate counts

Two HH001 listing submissions can produce one current HH001 dataset row if the link upserts on household ID. A filter may exclude a refusal. A long repeat can produce several destination rows from one submission. Therefore, a simple equality between submission count and dataset row count is not a universal QA rule. Write the expected relationship for each link, then compare the actual output with it.

Source eventRaw submissionsExpected current registry rows
HH001 initial listing11
HH001 corrected listing21 if joined by HH001
HH002 listing32

Review data with actionable queries

Monitor incoming submissions for missing IDs, duplicate IDs, unexpected values, unusual duration, invalid combinations, and field-version differences. A query should identify the record, field, observed value, expected relationship, and question for the field team. โ€œBad dataโ€ is not a query.

Example QA query

Submission KEY 7fโ€ฆ, HH003, listing version 20260925: household size is 3 but the member repeat has two rows. Please confirm whether a member row was omitted or the size was entered incorrectly. Do not edit the source until the field team resolves it; keep the original export.

When reviewing a published dataset, add the publishing-link version and refresh time to the query. A stale attached copy can show an older head name even when the current server dataset is correct.

Export with provenance

For each test export, record server, form or dataset ID, export time, form versions included, filters, and format. Keep an unmodified raw copy and do analysis in a separate file. If the form has repeats, inspect the repeat-level output and join keys. If datasets contain current state, keep a dated snapshot alongside source submission history so later investigators can reconstruct how a row changed.

Practice 22 ยท daily QA

Reconcile three events

  1. Use the table above to state expected raw and dataset counts.
  2. Write one query for a missing member row and one for an HH001 row that failed to update.
  3. Describe which evidence distinguishes a publishing failure from a stale device attachment.
  4. Make an export log entry with form ID, date, version, row count, and reviewer.
Expected reasoning

Two HH001 submissions remain in the raw data while one joined registry row shows the current value. A missing update requires checking the link; a stale device copy requires an attachment refresh.

Review status can delay publishing

SurveyCTO can hold incoming submissions for review instead of releasing them immediately. When that workflow is enabled, a submission may be received on the server yet absent from the dataset because it is awaiting approval. Approved submissions are released downstream; corrections made in review are applied to the published result. Record the form's review settings in your publishing contract. Test one record that is held and one approved record rather than assuming server receipt alone triggers the link.

Observed stateLikely next check
Submission received, dataset row absentIs it awaiting review? Also inspect link and filter
Submission approved, row still absentPublishing status, mapped fields, key, refresh
Corrected field differs from original entryCorrection log and approved published value

For a pilot, decide who may review, correct, approve, or reject. A data-quality correction can change a calculation or routing field, so rerun the downstream dataset and case checks after approval. Keep the raw record, correction history, and export settings together.

Knowledge check

Module 22 check

Attempts remaining: 3
1. What should a QA query identify?
2. Why compare submission counts and dataset rows?
Module 23 ยท Part VI ยท Operations

Protect and maintain a live workflow

Permissions, encryption, updates, and recovery

By the end of this module, you can:
  • Plan a safe form update
  • Recognize publishable data limits
  • Write a recoverable incident note

A connected form change has several dependents

Before changing a deployed question, search for its use in calculations, relevance, choice filters, dataset mappings, case routing, exports, and analysis scripts. A renamed hhid field can break a publishing join and a follow-up lookup. A changed stored outcome code can alter a 1/0 filter. A changed form_id can make a case's formids point to the wrong form. Keep a dependency map beside each maintained XLSForm.

ChangeCheck
Rename a fieldExpressions, publishing mappings, preloads, export scripts
Change a choice valueFilters, relevance, historical data meaning
Change form IDCases, attachments, links, device delivery
Change dataset keyUnique ID, joins, pulldata, search, case identity

Limit what can be published

SurveyCTO supports end-to-end encryption. Form data that remains end-to-end encrypted cannot simply be published into an ordinary server dataset. Fields deliberately marked publishable are an exception for values suitable to share in a connected workflow. Decide which fields are necessary for case routing or lookup, and avoid publishing a respondent's full answer set just because Add all is convenient. On a real project, document this choice with the data owner and test the effective export and dataset access.

Write a recoverable incident note

When a linked workflow fails, freeze the facts before changing settings: server URL, affected form and dataset IDs, form versions, device or browser, user role, case ID, submission KEY, expected state, observed state, and time. Keep a copy of the relevant dataset row and publishing-link configuration. Change one cause, rerun the same synthetic test, and record whether it resolved the issue. If local records are unsent, preserve them; do not clear app data as a first step.

Practice 23 ยท release review

Review a risky change

  1. Imagine renaming hhid to household_number in the listing form. List every downstream link that must be reviewed.
  2. Propose a safer path if only the respondent-facing wording needs improvement.
  3. Write a six-line incident note for an HH003 follow-up that shows no head name despite a received listing submission.
Model answer

Keep the stable variable name and change only the label when possible. For a missing name, inspect publishing, dataset row, attached copy, lookup key, and role/device context before editing the form.

Keep a change log that reaches beyond the workbook

For a live connected form, the release record should include the source workbook hash or filename, old and new form version, calculation tests, publishing-link changes, dataset schema changes, affected case queues, device update instruction, and rollback source. If the change is only to a question label, confirm that the variable name and stored choice values remain stable. If a mapped field must be retired, plan how old records will be exported and how the dataset column will be maintained.

Rollback is not merely uploading the old XLSForm. A publishing link, dataset row, or case assignment may already have changed. The recovery plan should say how to identify affected submission KEYs, whether a dataset snapshot exists, and who can approve a corrective republish or manual update. Test that plan on synthetic records before a production incident.

Knowledge check

Module 23 check

Attempts remaining: 3
1. What should be tested before changing a connected form ID?
2. Can an end-to-end encrypted value publish without being made publishable?
Module 24 ยท Part VII ยท Capstone

Capstone: household follow-up

Two forms, one lookup dataset, and a case queue

By the end of this module, you can:
  • Build the two forms
  • Test a publish-and-attach path
  • Document a field-ready release

Deliverable specification

Your final practice system has two forms, one household lookup dataset, and one cases dataset. The listing form records hhid, head name, village, consent, a repeat of members and ages, and adult_count. The follow-up form uses the case ID or entered household ID to pull the head name and records visit outcome, next visit date when needed, and a calculated publish flag. The registry has one current row per household. The cases dataset offers the follow-up form to a collector. Use only synthetic IDs HH001โ€“HH004.

ComponentTraining IDKey
Listing formhh_listing_traininghhid
Follow-up formhh_followup_trainingcaseid or hhid
Registry datasethh_registry_traininghhid_key
Cases datasetcases or a training-specific cases IDid

Part A: build and calculate

  1. Copy the listing starter. Complete its member repeat and adult-count calculation without reading the completed example first.
  2. Run the calculation test: ages 17, 18, and 65 give adult count 2; correct 18 to 16 and expect 1; remove the 65-year-old and expect 0.
  3. Run consent No and confirm roster values do not enter a completed refusal record unexpectedly.
  4. Compare your workbook with the completed example. Explain any difference in missing-value handling.

Keep a test matrix. An expression that compiles but produces the wrong value after an edit has not passed.

Part B: publish and preload

  1. Deploy the two forms on a designated practice server. Create the registry dataset with hhid_key as unique ID.
  2. Publish listing fields into the registry in wide format. Join hhid to hhid_key and record whether you backfilled existing submissions.
  3. Submit HH001 and verify server receipt and the dataset row.
  4. Attach the registry while uploading the follow-up form, then deploy it. Open HH001 and confirm that its head name is pulled from the right row.
  5. Correct HH001's head name in a second listing submission. Check the registry and then refresh the follow-up attachment. Record when the corrected name appears.
  6. Test HH999. The form should follow your defined missing-ID procedure.

Part C: cases and field access

  1. Prepare the cases CSV with HH001, HH002, and HH003. Set formids=hh_followup_training.
  2. In a training cases dataset, load the CSV and enable Manage Cases on the intended device if needed.
  3. Open HH001 through Manage Cases. Confirm the caseid stored in the follow-up is HH001.
  4. If test accounts are available, assign HH001 and HH002 to different users and verify both views. Do not use a real enumerator account merely to make the lab pass.
  5. Submit a callback for HH001. If you configure a case-state publishing link, verify the row's new state and its visibility after refresh.

Part D: negative and recovery tests

Failure injectedWhat to inspect first
HH001 exists in listing submissions but not registryPublishing link, filter, field relevance, key, and refresh
HH001 exists in registry but name is old on deviceAttached copy and device refresh
Collector cannot see HH002Cases dataset, role dataset ID, users/roles filter, refresh
One member disappears after reducing repeat countSaved form and roster-edit procedure
Late callback arrives after completionWhether Replace reverses case state

For each failure, change one setting or source value, rerun the same synthetic case, and write the result. Preserve original submissions and the prior workbook.

Capstone evidence pack

Hand the system to another analyst

  1. Maintained listing and follow-up XLSForms with versions.
  2. Registry schema, cases CSV, and each Publish into field map.
  3. Calculation test table and end-to-end HH001 trace: submission KEY โ†’ registry row โ†’ follow-up preload โ†’ case-managed submission.
  4. Role visibility results, offline/device test, and a negative-test log.
  5. One-page release note naming the server, owners, refresh procedure, open issues, and rollback source files.
Self-review

Can another analyst reproduce the adult count after an edit? Can they tell whether HH001 is absent because of publishing, attachment refresh, or visibility? Can they reconstruct the latest dataset row from source submissions? If any answer is no, improve the evidence pack.

Knowledge check

Module 24 check

Attempts remaining: 3
1. Which evidence shows the two-form connection worked?
2. What should a capstone handoff identify?
Module 25 ยท Part VIII ยท Questionnaire practical

Practical 2: sample and case setup

Prepare the ten-household sample and resolve missing inputs

By the end of this module, you can:
  • Audit the supplied case list
  • Set the preload and cases contracts
  • Test consent and case ID handling

Bring your build-along together

You have a household roster and education questionnaire and a ten-row household sample. Continue the form you began during the earlier build-along milestones: make it case-managed, test it with invented answers, and prepare its main and repeat exports for review. The HH001โ€“HH004 connected-data example is a separate teaching case. Keep the questionnaire beside the workbook; its question numbers are the specification.

Inspect the source before programming

Source itemWhat the file suppliesProgramming decision
caseidHH-2024-001 through HH-2024-010Keep as text. It joins Manage Cases, preload, and export.
LocationProvince, municipality, barangay, purokPreload by exact caseid.
Barangay IDAbsentThe derived training preload adds B001โ€“B010 only for practice. A real study needs an approved ID mapping.
PeopleHousehold head namesUse as a confirmation label, never as a unique key.
Assignmenttreatment_armKeep in the preload for analysis; decide explicitly whether collectors should see it.
StaffNo enumerator or supervisor master listThe reference form asks for names and IDs. Do not describe these fields as automatically populated.

Open the two CSVs and check all ten IDs, duplicate IDs, blank location fields, and unexpected leading or trailing spaces. The derived B-codes are a teaching device, not authoritative geographic codes.

Create the case and preload contracts

The included practice_cases.csv uses id,label,formids,users,roles,sortby,enumerators. Its id equals the source caseid; its formids is hh_roster_education_training. Blank user, role, and enumerator values make the training rows visible to users of that case list. Before assigning live work, replace them with approved test identities and verify each user's effective view.

Create a server dataset from practice_household_preload.csv, set caseid as its unique ID, and attach it when uploading the form. The reference form uses pulldata('practice_household_preload', 'province', 'caseid', ${hhid}) and the same pattern for other fields. Keep the dataset ID, attachment name, and key identical. Upload practice_cases.csv to the training cases dataset. The form should also allow a manual ID when opened outside Manage Cases; that path is for testing, not for replacing case assignment.

Consent is a real branch

Q0.1 asks whether the respondent consents. Code Yes and No as stored values, calculate a refusal flag, and put the interview inside a group relevant only when consent is Yes. Test both paths. A No response should create a refusal submission with case and staff identification, while the roster and education fields remain empty. Do not force a respondent into the interview to make the form complete.

Practical 2A ยท setup

Prepare two cases

  1. Use HH-2024-001 to verify the preload shows Bruno Fernandez in Iloilo, Pavia, Mabini.
  2. Use HH-2024-002 to verify a different location and a control assignment. Check that the assignment is stored without announcing it in the interview.
  3. Try HH-2024-999. Record the exact missing-ID message and confirm interview fields are unavailable.
  4. Answer consent No for one valid case; inspect the submission and confirm the refusal flag is 1.
Evidence to retain

Keep the case row, preload row, form version, test ID, displayed confirmation, and refusal export. A visible case label alone does not prove that pulldata() found the correct row.

Knowledge check

Module 25 check

Attempts remaining: 3
1. The source household list has no barangay ID. What should the programmer do?
2. Which field joins a case row to the household preload?
Module 26 ยท Part VIII ยท Questionnaire practical

Practical 2: program the questionnaire

Roster, partial ages, phones, and education

By the end of this module, you can:
  • Implement three linked repeats
  • Apply the questionnaire's skip patterns
  • Identify decisions the source leaves open

Use the PDF as a field specification

Program from the PDF first, then compare with the reference. The reference is an executable teaching adaptation, not an approved production instrument. It asks household size before the name repeat to control its count; the paper questionnaire instead ends name collection through Q1.7. Record this change in your specification and test how a collector corrects the count.

Part A: names and confirmation

Paper questionReference field or ruleAcceptance test
Q1.1โ€“Q1.3Names inside members repeat; member_index=index()Member 1 is the head; repeat exactly the reported size.
Q1.4โ€“Q1.5aSuffix and nickname follow-ups use relevance on Yes or Other.Neither follow-up appears after No, unknown, or refusal.
Q1.6Relationship constraint allows code 1 only for member 1.Second head is rejected before roster confirmation.
Confirmationjoin(', ', ${full_name}) lists names; count-if() counts spouses.Two spouses or a rejected roster blocks progress until corrected.

Do not assume a name is unique. Member order is the link between the name, demographic, and education repeats. Changing repeat count after data have been entered needs a documented correction procedure.

Part B: demographics and age

The demographics repeat has repeat_count=count(${members}). Each iteration uses indexed-repeat(${full_name}, ${members}, index()) so the name beside Q1.8โ€“Q1.12 comes from the matching roster position. Test a household with a child first, an adult second, and another child third; a skipped member must not shift the later name.

The source allows -999 (unknown) and -888 (refusal) for birth day, month, and year. The reference calculates completed years from the three components when all are positive. If any component is unavailable, it asks Q1.10 for age and unit; days and months become zero completed years, while years use the reported number. The reference bounds days at 29, months at 11, and years at 110. Q1.12 appears only above age 13; Q1.12a appears for codes 1โ€“5. Test ages 13 and 14 and a birthday immediately before or after the interview date.

The PDF has no Q1.11, and its partial-date specification does not state how to reject impossible calendar dates such as 31 February. The reference checks component ranges but leaves cross-component date validation unresolved. Decide the interview rule with the study owner before field release. Record whether a partial birth date should produce an estimated age, an unknown age, or an explicit follow-up.

Part C: phones and education

The paper form asks Q1.13 to select phone holders, then repeats Q1.14โ€“Q1.18 for up to three holders. The reference asks has_phone inside each demographic repeat and calculates phone_holder_count. It displays a warning above three but does not silently drop a fourth phone owner. The three-person cap requires a study decision: hard stop, first three by a documented rule, or a revised questionnaire. Mobile numbers are stored as text so an initial zero survives; the constraint accepts 11 digits starting with 0 and the explicit -999/-888 codes.

The education repeat also uses the full roster count. It pulls the name and calculated age at the same index, then puts Q2.1โ€“Q2.6 in a group relevant for age 5 or older. Using only the count of eligible members would mismatch names and ages when an under-five appears early. Q2.2 and Q2.3 appear for current attendees. Q2.4 appears only for non-attendees aged 5โ€“24. Q2.5 appears for every eligible member; Q2.6 excludes codes 00 and 99. The reference uses separate current-grade and highest-grade lists because the source adds 00 and 60 only for Q2.5.

Practical 2B ยท program and compare

Build the three-repeat form

  1. Complete the starter workbook using the paper questions. Keep the exact field names in your own mapping sheet.
  2. Test a four-person roster with ages 4, 5, 13, and 25. Explain which education and marital questions appear for each member.
  3. Change the age-five member to age four. Confirm the education answer becomes irrelevant without attaching that person's answers to another member.
  4. Test 91234567890, 09123456789, -999, and -888 as phone responses. Use invented numbers only.
  5. Compare your workbook with the reference and list every deliberate deviation from the paper instrument.
Answer check

Education is skipped for age 4 and shown for ages 5, 13, and 25. Marital status is shown only above 13. A non-attending 25-year-old skips Q2.4. The 11-digit number beginning with 0 passes; the one without 0 fails. Repeating the full roster count preserves member indexes.

Knowledge check

Module 26 check

Attempts remaining: 3
1. Why does the education repeat use the full roster count and then skip children under five?
2. What should Q2.4 do for a non-attending 25-year-old?
3. How should 00 and 60 be used in the grade lists?
Module 27 ยท Part VIII ยท Questionnaire practical

Practical 2: release test

Run boundary tests, inspect exports, and document a release

By the end of this module, you can:
  • Execute a case-by-case test matrix
  • Inspect main and repeat exports
  • Decide whether the form is ready for field use

Run a release test, not just a preview

The supplied matrix contains twelve tests. Add the observed result, form version, device or browser, and a pass or fail for each. A form that compiles has only passed a syntax check. Test through Manage Cases on a device or an enabled web-collection account, then compare with the designer's Test view.

Test sequences that can expose a wrong calculation

SequenceExpected behaviorWhy it matters
HH-2024-001, correct preload, then unknown IDFirst shows the right household; second stops before consent.A case label does not prove a successful preload.
Consent No, then Yes on a separate testRefusal has no roster; consented case proceeds.Skipped answers must not survive a changed branch.
Age 4 โ†’ 5 โ†’ 4 in edit modeEducation appears, then becomes irrelevant again.Derived eligibility must update after correction.
Non-attending age 24 โ†’ 25Q2.4 disappears at 25.The upper age bound is inclusive at 24.
Two spouses; second headRoster cannot be confirmed until corrected.Relationship rules protect roster structure.
Impossible date; four phone ownersRaise open issues in the release log.The reference intentionally does not resolve these protocol decisions.

Inspect the exports

Submit only synthetic cases. Export the main submission data and the member, demographic, and education repeats. Join repeat rows by their parent submission key and repeat index; do not assume that a member name is a join key. Check that a refusal has no member rows; that each consented household has the expected number of member and demographic rows; and that education rows for under-fives have no Q2 answers. Preserve the original submissions and an export copy when correcting a mistake.

If you publish a household-level summary into a server dataset, map hhid to a unique caseid column, and publish only non-sensitive indicators needed for follow-up, such as roster count or an interview status. Do not publish names, phone numbers, or treatment assignment into a broad case list. Define whether a correction replaces a prior summary row and test a second submission for the same case.

Release decision

Before calling this field-ready, obtain decisions on the missing barangay ID source, staff roster and auto-fill, the Q1.7 roster-length adaptation, impossible dates, partial age, and the three-phone rule. Confirm consent wording with the study owner. Then run the matrix again on the approved form version and target device. The included reference workbook remains a training aid until those decisions are implemented and tested.

Practical 2C ยท handoff

Give another analyst a reproducible package

  1. Provide the maintained XLSForm, its version, the approved questionnaire, the case and preload CSVs, and a field mapping.
  2. Provide completed tests P01โ€“P12, including observed screenshots or written evidence and open issues.
  3. Show one consented case and one refusal in the exports, plus the repeated member records.
  4. State who owns case assignment, preload refresh, form deployment, QA review, and correction approval.
Completion standard

An analyst who did not program the form should be able to reproduce the relevant and skipped questions for an under-five, a 24-year-old, a 25-year-old, a refusal, and an unknown case. Any unresolved protocol decision remains an open release issue.

Knowledge check

Module 27 check

Attempts remaining: 3
1. What does the reference form do with an impossible date such as 31 February?
2. Where should repeated member and education responses be checked?