Data Resources
Stanford Electronic Health Records
Stanford EHR data includes rich multimodal ophthalmic data on over 300,000 ophthalmology patients. EHR data is updated yearly and includes standard structured fields (demographics, billing and procedure codes, medications, lab findings), eye exam findings (VA, IOP, refraction, and others). Free-text clinical progress notes are available as part of EHR data. Visual field data are available (Humphrey) as well as imaging data (many types, including photography, OCT, and others - but imaging last update was mid-2022).
- How do researchers gain access to Stanford EHR data?
- To gain access to fully-identified Stanford EHR data, you will need an IRB covering your project.
- If you just want to do some initial counts of patients to determine feasibility, I highly recommend Stanford Cohort Discovery Tool, which allow you to do initial de-identified searches without IRB.
- If you want EHR data but DO NOT need eye-specific data, then you can actually get this directly through STARR tools, using Stanford Cohort Discovery, STARR Chart Review, and STARR Data Delivery.
- If you need eye-specific data, such as visual acuity, IOP, imaging, fields, etc. we are happy to help you. Our staff must be on your IRB for your project. It is best if you already have a list of MRNs you need this eye-specific data for. Contact us for more information.
- What computational tools do I need to work with Stanford EHR data?
- Fully-identified and therefore high-risk data with PHI should be accessed within computing environments that are approved for PHI. This includes, for example, Stanford Nero-GCP.
Sight Outcomes Research Collaborative (SOURCE)
SOURCE (sourcecollaborative.org) a consortium of academic ophthalmology departments sharing de-identified electronic health record and ocular diagnostic test data on eye care recipients. Stanford is a member and contributes our EHR structured data, free-text notes, and visual field data. The data is aggregated, cleaned, and harmonized in a repository that faculty and trainees from participating sites can tap into and use for research and quality improvement projects.
- How do researchers gain access to SOURCE data?
- To obtain access to SOURCE data for a research or Q/I project, a researcher must be at an institution who is actively contributing data to the consortium. The researcher submits a brief proposal describing the proposed project, the variables needed, etc. All proposals get reviewed by the SOURCE Research Committee to ensure that the project is feasible and doable and not being done for nefarious purposes. Once the proposal is approved, a member of the SOURCE Data Office will extract the necessary data elements for the project and put them in a virtual sandbox on a HIPAA-compliant server at the University of Michigan. The virtual sandbox will also have statistical software to permit the researcher to carry out their analyses. The researcher will VPN into the virtual sandbox to carry out their analyses. To protect the data, no data can be downloaded or removed from the virtual sandbox by the researcher. Once all analyses are completed, a member of the SOURCE Data Office can pull the results off the sandbox and provide them to the researcher for presentations and publications.
- Where do I find the proposal form I need to fill out to access SOURCE data?
- https://www.sourcecollaborative.org/tools (first link)
- How long does it take to get access to SOURCE data?
- Proposals are reviewed and approved by the SOURCE research committee roughly once a month. Your proposal might be approved or recommended to be revised and resubmitted. Once your proposal is approved, it will go into a queue so that the SOURCE data team can extract your requested data. The queue is of variable length but probably weeks+ at this time.
- What data types can be requested from SOURCE?
- De-identified structured EHR data are available to all. This includes patient demographics, diagnoses, procedures, eye exam information, lab measurements, surgery information including supplies and implants used and others.
- Because Stanford submits EHR free-text clinical progress notes, we can also request access to medical concepts derived from progress notes from other SOURCE sites but currently this is just UMich notes.
- Because Stanford submits Humphrey Visual Fields, we can also request access to HVF from other SOURCE sites (this includes UMich and a few other sites)
- Individual-level social determinants of health is also available (income, education, etc.) but this is on a subset of patients and it costs extra to access. Please discuss with Dr Wang (and eventually, with Dr Stein, Chief Data Officer of SOURCE).
- Community-level social determinants of health (based on zip code) are available with no extra cost for most patients, including urban/rural commuting codes (RUCA) and Distressed Communities Index (DCI). A subset have ADI national rank also.
- What software is available for data analysis in SOURCE?
- Standard statistical software is available, such as SAS, SPSS, R, Python...
- Who is eligible to apply for SOURCE data access?
- Faculty at member institutions, including Stanford.
- What data are not available from SOURCE?
- SOURCE does not contain much data on inpatient care, e.g. hospitalizations, IV medications, and probably not ED visits either.
- If a patient wasn't seen by an eye care provider, they are likely not in the SOURCE database. This makes this database not ideal if your study question is about incidence of some eye condition as a complication of a systemic condition, because there's no way to capture the denominator, e.g. "all patients with some systemic condition" since if they never saw eye care provider, they're not in the SOURCE database.
- Location-based data are not available in SOURCE, to protect anonymity. We do not know the geographic area (east coast, west coast, midwest). We also don't know such details as number of miles from patient address to medical center. The only thing we really know about location is the rural-urban commuting codes (RUCA).
Other Available Datasets
- All of Us (https://researchallofus.org/): A US national database of EHR as well as patient-reported surveys on SDoH, lifestyle, lots of other factors. Check out the website for full details on what data are available!
- Available to researchers from Stanford.
- No eye exam data, but does include systemic exam data such as blood pressure, HR, etc.
- You must request approval for access: https://redcap.pmi-ops.org/surveys/?s=TFF9RN7NLLHC9T4N
- A PI can fill out one request form (and individual lab members can be listed). It will eventually be approved and routed to Stanford Research Related Agreements.
- After access granted, visit the researchallofus.org website to create a researcher account. Complete the data access training. Choose the tier of access you need.
- Every new user has $300 of computing credits - enough to do a standard research study or two.
- IRIS Registry: National registry of EHR data submitted by ophthalmology practices. Includes eye exam data. We have access through Stanford.
- Cosmos: Epic's EHR database.
- Please see this website for more detail about access through Stanford: https://med.stanford.edu/starr-tools/other-resources/cosmos.html
- Includes visual acuity, refraction, eye pressure measurements.
- Data access is complicated and can take quite a long time.
- TriNetX: EHR and insurance claims database. Stanford access has expired.
- Truven MarketScan: Very large insurance claims database hosted through Stanford PHS.
- See Stanford PHS website for more details: https://stanfordphs.redivis.com/datasets
- Access to this dataset is sponsored by the department of ophthalmology so no extra cost for PI's.
- Claims only: no eye exam data.
General Getting Started
Have a research question (or an inkling of one) and wondering where to get started? Definitely check out the Practical Foundations of Data Science in Ophthalmology series. Register to get access to all the lectures and materials. Session 2 of that series goes over these datasets (and more!) to give you an idea of where to get started.
- Which dataset is right for my research question?
- Please see handy flowchart here for an idea...(click on the image the see the entire flowchart)
- A summary table is also below here.
| Dataset | Type | Size | Laterality? | Eye Exam? | Text? | Imaging? | Systemic Data? | SDoH? |
| Stanford EHR | EHR | Medium | Yes | Yes | Yes | Yes | Some | No |
| SOURCE | EHR | Large | Yes | Yes | Some, soon | Some HVF | Decent | For a fee |
| Truven | Claims | Large | No | No | No | No | Yes | No |
| All of Us | EHR | Large | No | No | No | No | Yes | Yes |
| IRIS | EHR | V Large | Yes | Yes | No | No* (some OCT #s) | Some | Some |
| Medicare | Claims | V Large | Some | Yes | No | No | Yes | No* (linked data) |
| Cosmos | Claims+EHR | V Large | No | No | No | No | Yes | No |
| TriNetX | Claims+EHR | V Large | No | No | No | No | Yes | No |