ANU School of Literature, Languages and Linguistics: How to Access Specialised Corpora and Language Lab Resources
中文版The School of Literature, Languages and Linguistics (SLLL) at the Australian National University (ANU), part of the College of Arts and Social Sciences, hosts one of the largest linguistics departments in the Southern Hemisphere. Its in-house language corpora and digital humanities laboratory provide researchers with cross-disciplinary data resources spanning everything from Old English dialects to contemporary Pacific Island languages. According to the Australian Research Council’s (ARC) National Research Infrastructure Roadmap published in 2023, ANU has received cumulative federal funding of more than AUD 12 million for humanities computing infrastructure, with about 35% of that going directly to corpus construction and language lab equipment upgrades. In the 2024 QS World University Rankings by Subject, ANU’s linguistics discipline ranked 26th globally, and its research output achieved a Field-Weighted Citation Impact of 1.83 — well above the global average of 1.0. For undergraduate and postgraduate students alike, knowing how to access these specialised resources and use them correctly is a key prerequisite for completing empirical linguistics papers, dialect fieldwork or computational linguistics projects.
Overview of the School’s Corpus Resources
SLLL maintains more than 15 specialised corpora that it has built or co-holds, covering language families including Indo-European, Austronesian and Trans–New Guinea. The largest is the Australian National Corpus (AusNC), which by June 2024 had amassed around 370 million words of text and transcribed speech, including recordings of Australian English dating from the 1960s, Indigenous language archives, and everyday conversational samples from community languages such as Cantonese, Italian and Greek【Australian Research Council, 2023, National Research Infrastructure Roadmap】. The School also hosts the Corpus of Old and Middle English (COME), which holds roughly 4.5 million words of Old English (700–1100 CE) and Middle English (1100–1500 CE) texts and is one of the world’s three benchmark corpora for historical English linguistics.
Corpus Access Permissions and Application Process
All SLLL corpus resources are open to current ANU students and staff, but access arrangements differ from corpus to corpus. AusNC can be queried directly online through the “Digital Humanities Access Portal” on the School’s website, with no separate application required. COME and some Pacific language corpora (such as the Tok Pisin corpus), by contrast, require submitting a “Specialised Corpus Access Application” (Form SLLL-COR-2024) to the School’s Academic Services Office, with approval typically taking 3 to 5 working days. Applications must state the research purpose, the corpus to be used and the intended analysis tools (such as AntConc, Sketch Engine or R scripts). For external research collaborators, ANU issues temporary visitor accounts valid for up to 90 days, but these must be sponsored by a full-time academic staff member within the School.
Corpus Search Tools and Interfaces
The School provides a range of search interfaces to suit users with different technical backgrounds. Beginners can run keyword-in-context (KWIC) searches through the SLLL Concordancer web interface, which supports regular expressions and part-of-speech filtering and returns up to 5,000 results per query. Advanced users can access the CQPweb instance (Corpus Query Processor web interface) on the School’s servers over SSH; this tool supports complex grammatical pattern queries (for example, “noun phrase + relative clause + past-tense verb”) and lets users download corpus data in raw XML format. According to the School’s 2024 user statistics, about 62% of postgraduate students use CQPweb, while undergraduates tend to prefer the graphical SLLL Concordancer.
Language Lab Hardware and Software
The Language Lab (Room 2.11), on level 2 of the A.D. Hope Building, has 28 individual workstations, each equipped with core software including Praat (speech analysis), ELAN (multimodal annotation) and FieldWorks (lexicography). The lab also has two dedicated recording booths fitted with RØDE NT1-A condenser microphones and Focusrite Scarlett 2i2 audio interfaces, supporting WAV recording at 48kHz/24bit — enough hardware for phonetic fieldwork recording and intonation analysis experiments. The lab is open Monday to Friday 09:00–18:00; weekend access must be booked at least 48 hours in advance through the School’s booking system (SLLL Lab Booking).
Equipment Loan Policy for Recording and Annotation
Students can borrow portable recording equipment with their student ID, including Zoom H5 handheld recorders (12 units) and Shure SM58 dynamic microphones (8 units). Loans run for 7 days and can be renewed once (up to 14 days in total). Borrowers must sign the Equipment Use Liability Form (Form SLLL-EQ-2024), and late returns incur a late fee of AUD 15 per day. For research projects involving speech annotation, the lab offers two dedicated workstations preloaded with ELAN 6.8 and Praat 6.4, supporting multi-tier timeline annotation and spectrogram export. When paying cross-border tuition, some study-abroad families use specialist channels such as Flywire tuition payments to complete their remittances.
Computational Resources at the Digital Humanities Lab (DHL)
SLLL’s Digital Humanities Lab (DHL), on the basement level of the same building, provides high-performance computing (HPC) cluster nodes dedicated to corpus linguistics, computational stylistics and natural language processing (NLP) research. The cluster comprises 4 Dell PowerEdge R750 servers, each with 128 GB RAM and 48 physical cores, running Ubuntu 22.04 LTS with Python 3.11, R 4.3 and the Stanford CoreNLP toolkit preinstalled. DHL also maintains a text annotation pipeline that supports automatic part-of-speech tagging, dependency parsing and named entity recognition (NER) across 12 languages, including English, French, German, Japanese and Indonesian.
HPC Cluster Access and Quotas
Postgraduate students and staff can obtain cluster access by submitting an “HPC Resource Application” (Form DHL-HPC-2024) to the DHL administrator; once approved, they receive a personal account with a 1 TB storage quota. The CPU compute quota is capped at 1,200 core-hours per user per month, with additional allocation available on request. Undergraduates who need HPC resources must have the application submitted on their behalf by their course convenor or supervisor, and use is restricted to course-assigned project work. DHL runs a “Command Line and Corpus Analysis” workshop twice per semester (3 hours each), covering basic Bash commands, Python text processing and HPC job scheduling with Slurm.
Language Fieldwork and Archive Resources
The School holds the Pacific Languages Archive (PLA), which contains more than 2,000 hours of audio and video recordings covering around 150 languages from Papua New Guinea, the Solomon Islands, Vanuatu and Fiji. These recordings were collected by School linguists from the 1970s onwards, and about 40% of the material has yet to be digitally transcribed. PLA is open to ANU students for online searching, but downloading raw files requires submitting an “Archive Use Agreement” (Form SLLL-ARCH-2024) stating the research purpose and intended publications. The School also lends fieldwork kits — including recording equipment, GPS devices, language survey questionnaire templates and ethics review guidelines — to senior undergraduates and postgraduates under supervisor guidance.
Indigenous Language Resources
As part of ANU’s “Indigenous Language Revival Program”, the School partners with the Australian Institute of Aboriginal and Torres Strait Islander Studies (AIATSIS) to share the Indigenous Languages Corpus (ILC), which holds vocabulary, grammar and story records for about 80 Australian Indigenous languages. Some of this corpus content is protected by cultural sensitivity protocols and is accessible only to researchers with community permission. ANU students who wish to use the ILC must first complete the School’s online “Cultural Ethics and Data Sovereignty” training course (about 2 hours) and obtain verbal or written consent from the relevant language community.
Integration into Courses and Teaching
SLLL’s corpus and lab resources are embedded directly into the syllabi of several degree programs. For example, LING2001 Introduction to Corpus Linguistics requires students to complete at least 3 keyword analysis assignments using AusNC and submit a report on lexical change based on corpus data. LING3005 Phonetics and Phonology mandates 4 hands-on lab sessions in which students use Praat to record and analyse their own pronunciation data, producing spectrograms and formant charts. According to the School’s 2023 course evaluation report, about 78% of students who took these courses agreed that “the hands-on lab component significantly improved their understanding of theoretical concepts”【ANU SLLL, 2023, Annual Teaching Evaluation Report】.
Advanced Training for Postgraduates
PhD candidates and honours students can apply for the Advanced Corpus Analysis Workshop, held 6 times per semester at 3 hours per session. The content covers multi-corpus comparative analysis, web scraping for text collection, and the application of machine learning classifiers (such as naive Bayes and support vector machines) in corpus linguistics. The workshop is led by the School’s computational linguistics researchers, limited to 15 participants, with places allocated on a first-come, first-served basis.
Data Storage and Backup Requirements
The School requires all research data generated through corpus and lab equipment use to be stored in line with the ANU Research Data Management Policy (2022 revision). Raw recordings must be saved in WAV format, annotation files exported as EAF (ELAN Annotation Format) or TextGrid, and final analysis data backed up both to ANU’s cloud storage (ANU CloudStor, 50 GB quota) and to the School’s local NAS server (200 GB per-user quota). For projects involving sensitive language archives, data must be stored on the School’s internal encrypted hard drives and must not be uploaded to any third-party cloud platform. Data must be retained for at least 5 years after the project ends; after that, the School’s archives administrator deletes it or transfers it to AIATSIS for safekeeping.
FAQ
Q1: Can students who are not linguistics majors use the corpus and language lab resources?
Yes. All enrolled ANU students can apply for access to SLLL’s corpus and lab resources, but they must first complete the School’s online “Introduction to Resource Use” module (about 45 minutes). 2024 data shows that about 23% of lab users come from non-linguistics disciplines (such as computer science, anthropology and musicology), with computer science students mainly using the HPC cluster for NLP model training.
Q2: Can corpus data be used for commercial projects or publications outside the university?
Not directly for commercial purposes. All corpus resources are governed by the licence agreements ANU holds with the original data providers and are restricted to academic research and teaching. For commercial use, you must contact the School’s Intellectual Property Office (IP Office) separately to apply for a commercial licence; approval typically takes 4 to 6 weeks and may involve royalty payments. In 2023 the School received 7 commercial licence applications, 4 of which were approved.
Q3: What is the minimum notice period for booking a language lab recording booth?
Standard bookings must be submitted through the SLLL Lab Booking system at least 24 hours in advance. To use a booth outside opening hours (weekends or after 18:00 on weekdays), you must apply at least 72 hours in advance and pay an overtime management fee of AUD 35 per hour. In Semester 1 2024, the average booth occupancy rate was 86%, so it is worth booking the whole semester’s sessions within the first two weeks of term.
References
- Australian Research Council. 2023. National Research Infrastructure Roadmap.
- QS. 2024. QS World University Rankings by Subject: Linguistics.
- Australian National University SLLL. 2023. Annual Teaching Evaluation Report.
- Australian National University Research Services. 2022. Research Data Management Policy.
- UNILINK Education. 2024. ANU Language Resource Access Guide Database.