← THE FIELD RECORD
TREQS ON A HOSPITAL DATABASE

A question in English, a query in SQL

19 June 2020 · Online, over Zoom. Presented from the Netherlands, with a co-presenter joining from Virginia Tech. · RESEARCH

An hour and three quarters on Zoom, spent on a model that turns a clinician's plain question into SQL against de-identified intensive care records, and a demo that ran.

Slide from the session showing the TREQS architecture: an encoder-decoder with attention that takes a natural language question and emits a SQL query, annotated in red during the talk. The paper citation sits along the bottom.

Nobody was in a room. On 19 June 2020 the session ran on Zoom for an hour and three quarters, and what survives of it is the screen capture: 2736 by 1824 pixels, one audio channel, a mouse pointer moving across slides. There are no photographs of this one because there was nothing to photograph. It opened on Tarry's Forbes column from the previous August, Software Ate The World, Now AI Is Eating Software, scrolled live in a browser with the dk mark sitting in the corner of the header art. Then the screen changed hands.

The subject was TREQS, a TRanslate-Edit Model for Question-to-SQL query, published at The Web Conference in April 2020 by Ping Wang, Tian Shi and Chandan K. Reddy. The problem it takes on is narrow and real. A clinician wants to know how many female patients underwent an abdomen artery incision. The answer sits in a relational database of demographics, lab tests, prescriptions, diagnoses and procedures, joined on subject and admission IDs, and getting it out means writing SQL. TREQS reads the English and writes the query, generating from a vocabulary distribution where it can and copying words straight out of the question where it must.

The data was MIMIC-III, 46,520 de-identified intensive care patients treated at Beth Israel Deaconess between 2001 and 2012. From that the authors built MIMICSQL on 100 randomly sampled hospital admissions, in two versions: template questions produced by machine, then the same questions rewritten by hand into natural language. The gap between the two versions is the whole difficulty. One asks for the number of patients whose gender is f and whose drug name is amitriptyline. The other asks, among patients treated with amitriptyline, for the count of female patients. Same SELECT statement, almost no shared vocabulary.

Later in the session the slides gave way to a repository. main.py open on GitHub, the argument list scrolling past with a beam size of five and a flag for copying words. Then a Flask server on 127.0.0.1 port 5000 with two controls, a question box and a button marked Generate Query. A local demo on a laptop, unstyled, and it ran.

The session did not stop at SQL. It carried on into a hierarchical attention retrieval model for healthcare question answering, the 2019 paper from the same lab, working through the shapes real patient questions take: what, how, yes or no, who. After that came an architecture for reading CT scans.

FROM THE FIELD · 09 FRAMES
Tarry Singh on camera during the Zoom session, wearing earphones.
Tarry Singh on camera during the Zoom session, wearing earphones.
The session opens on Tarry Singh's Forbes column, Software Ate The World, Now AI Is Eating Software, scrolled live in a browser. The dk mark sits in the corner of the header illustration.
The session opens on Tarry Singh's Forbes column, Software Ate The World, Now AI Is Eating Software, scrolled live in a browser. The dk mark sits in the corner of the header illustration.
Slide showing the relational structure of the medical records: demographic, lab test, prescription, diagnosis and procedure tables joined on subject and admission IDs.
Slide showing the relational structure of the medical records: demographic, lab test, prescription, diagnosis and procedure tables joined on subject and admission IDs.
Slide pairing four template questions with their natural language rewrites, marked up in red during the talk.
Slide pairing four template questions with their natural language rewrites, marked up in red during the talk.
The live demo. A local server on 127.0.0.1 port 5000 with a question box and a Generate Query button, the co-presenter's video tile at the right.
The live demo. A local server on 127.0.0.1 port 5000 with a question box and a Generate Query button, the co-presenter's video tile at the right.
The TREQS repository open on GitHub, showing the argument list in main.py.
The TREQS repository open on GitHub, showing the argument list in main.py.
The MIMICSQL documentation on GitHub, listing the demographic, diagnosis, procedure, prescription and lab fields and how the question set was built.
The MIMICSQL documentation on GitHub, listing the demographic, diagnosis, procedure, prescription and lab fields and how the question set was built.
Slide of example medical questions grouped by what, how, yes or no, and who, from the 2019 healthcare question answering paper.
Slide of example medical questions grouped by what, how, yes or no, and who, from the 2019 healthcare question answering paper.
MATERIALS
Session recordingTO CONFIRM

Zoom screen capture, 1 hour 44 minutes, 2736 x 1824. Held in the archive, not published.

Open repository, shown on screen during the session.

Source paper

Ping Wang, Tian Shi and Chandan K. Reddy, "Text-to-SQL Generation for Question Answering on Electronic Medical Records", Proceedings of The Web Conference (WWW), Taiwan, April 2020.

Second paper discussed

Ming Zhu, Aman Ahuja, Wei Wei and Chandan K. Reddy, "A Hierarchical Attention Retrieval Model for Healthcare Question Answering", Proceedings of The Web Conference (WWW), San Francisco, May 2019.

WITH
Virginia Tech
← EARLIERThree medical cases and no roomLATER →AI Transformation, Live to Camera