We are pleased to announce the 2nd Workshop on Sign Language Processing (WSLP 2026), co-located with EMNLP 2026, to be held in Budapest, Hungary, on October 29, 2026.
WSLP 2026 aims to provide a dedicated forum for researchers, practitioners, and industry professionals working on sign language technologies. The workshop will showcase recent advances in sign language understanding, translation, recognition, generation, and multimodal AI, with a special emphasis on low-resource and underrepresented sign languages.
In addition to paper presentations, the workshop will feature shared tasks, invited talks, and discussions on emerging challenges and opportunities in the field.
🔑 Important Links
🌐 Workshop Website:
https://exploration-lab.github.io/WSLP-2026/
📄 Submission Portal (OpenReview):
https://openreview.net/group?id=EMNLP/2026/Workshop/WSLP
🏆 Shared Tasks:
https://exploration-lab.github.io/WSLP-2026/task/
💬 Workshop Discord:
https://discord.gg/x6tHXEFyW7
💬 Shared Task Discord:
https://discord.gg/ty7dqTHC2v
🗓️ Key Dates (AoE)
* Paper Submission Deadline: August 24, 2026
* Notification of Acceptance: August 30, 2026
* Camera-Ready Papers Due: September 10, 2026
* Workshop Date: October 29, 2026
🔍 Topics of Interest (Including but Not Limited To)
* Low-resource and underrepresented sign languages
* Sign Language Recognition (SLR)
* Continuous and Gloss-free Sign Language Translation (SLT)
* Sign Language Generation and Avatars
* Multilingual, Cross-modal, and Zero-shot Learning for Sign Languages
* Large Language Models and Multimodal Models for Sign Languages
* Cultural and Dialectal Variation in Sign Languages
* Corpus Creation, Annotation, and Data Resources
* Fairness, Ethics, Accessibility, and Real-world Applications
🏆 Shared Tasks at WSLP 2026
* Indian Sign Language to English Translation
* Isolated Sign Recognition
* Word/Sign Presence Prediction
We particularly encourage submissions that focus on underrepresented sign languages, inclusive technologies, and real-world deployment challenges.
We warmly invite the community to submit their latest research and join us in advancing sign language technologies that foster accessibility, inclusion, and communication for all.
In this newsletter:
Fall 2026 LDC data scholarship program
New publications:
Parsed Early English Books Online -Text Creation Partnership<https://catalog.ldc.upenn.edu/LDC2026T09>
Multi-Language Conversational Telephone Speech 2014 - Slavic Group<https://catalog.ldc.upenn.edu/LDC2026S10>
LORELEI Arabic Representative Language Pack<https://catalog.ldc.upenn.edu/LDC2026T08>
________________________________
Fall 2026 LDC data scholarship program
Student applications for the Fall 2026 LDC data scholarship program are being accepted now through September 15, 2026. This program provides eligible students with no-cost access to LDC data. Students must complete an application consisting of a data use proposal and letter of support from their advisor. For application requirements and program rules, visit the LDC Data Scholarships<https://www.ldc.upenn.edu/language-resources/data/data-scholarships> page.
________________________________
New publications:
Parsed Early English Books Online - Text Creation Partnership<https://catalog.ldc.upenn.edu/LDC2026T09> (EEBO-TCP), developed by LDC, is a part-of-speech tagged and syntactically parsed version of the EEBO-TCP<https://www.textpartnership.net/pages/eebo-tcp.html> collection of Early Modern English texts. The corpus consists of 59,433 texts dating primarily from 1600-1700, with a smaller number of texts from earlier and later periods. It contains 48.6 million parsed sentences (trees) comprising more than 1.5 billion tokens. The parses were produced automatically using a parser trained on the Penn Helsinki Parsed Corpus of Early Modern English, part of the Penn Parsed Corpora of Historical English (LDC2020T16)<https://catalog.ldc.upenn.edu/LDC2020T16>. The automatically generated parses were not manually reviewed.
Parsed EEBO-TCP includes the CorpusSearch 2 program and associated documentation. This tool allows users to search the data for syntactic structure, word sequences, and words. An alternative version of CorpusSearch 2 for use on very large corpora is also included in this release.
2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.
*
Multi-Language Conversational Telephone Speech 2014 - Slavic Group<https://catalog.ldc.upenn.edu/LDC2026S10> was developed by LDC and is comprised of 20 hours of Polish and Russian telephone speech. The data was collected to support research and technology evaluation in automatic language identification; portions of these recordings were used in the NIST 2015 and 2017 language recognition evaluations<https://www.nist.gov/itl/iad/mltg/language-recognition>. The collection focused on language pair discrimination for 20 languages/dialects, some of which could be considered mutually intelligible or closely related.
This corpus contains 96 recordings: 61 Polish recordings and 35 Russian recordings. Participants were recruited by native speakers who contacted acquaintances in their social network. Those native speakers made one call, up to 8 minutes, to each acquaintance. Human auditors labeled the calls for language, callee gender, native speaker status, and speech clarity.
2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.
*
LORELEI Arabic Representative Language Pack<https://catalog.ldc.upenn.edu/LDC2026T08> contains 2.4 million words of Arabic monolingual text, 930,00 words of which were translated into English, and 225,000 Arabic words translated from English data. Nearly 88,000 words were annotated for simple named entities, and approximately 29,000 words were annotated for full entities (including nominals and pronouns). In addition, over 10,000 words were annotated for noun phrase chunking, 18,000 words were labeled with semantic annotation, and situation frame annotation was applied to more than 24,000 words.
The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages in the context of emergent situations. Representative languages were selected to provide broad typological coverage.
The knowledge base for entity linking annotation is available separately as LORELEI Entity Detection and Linking Knowledge Base (LDC2020T10)<https://catalog.ldc.upenn.edu/LDC2020T10>.
2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.
To unsubscribe from this newsletter, log in to your LDC account<https://catalog.ldc.upenn.edu/login> and uncheck the box next to "Receive Newsletter" under Account Options or contact LDC for assistance.
Membership Coordinator
Linguistic Data Consortium<ldc.upenn.edu>
University of Pennsylvania
T: +1-215-573-1275
E: ldc(a)ldc.upenn.edu<mailto:ldc@ldc.upenn.edu>
M: 3600 Market St. Suite 810
Philadelphia, PA 19104
Dear All,
We are excited to announce the PTA: From Pretrained Representations to Acting Agents workshop at NeurIPS 2026 in Sydney, Australia! 📋 We’re calling for submissions!
⏰ Paper submission deadline: Aug 29, 2026 (AoE)
🌐 Website: https://ptaworkshop.github.io/
📋 Submission link: https://openreview.net/group?id=NeurIPS.cc/2026/Workshop/PTA
Our goal is to explore how to obtain actionable pretrained representations and align them with control agents at test time—bridging pretraining and test-time decision making.
Call for Papers:
- Specific topics include, but are not limited to:
(1) What makes a pretrained representation actionable?
(2) How to align with and use pretrained representations at test time?
(3) How to adapt pretrained knowledge to new tasks?
(4) How to evaluate actionable representations?
(5) What can go wrong when pretrained agents act?
- We welcome both short papers (4 pages) and long papers (up to 9 pages), formatted using the NeurIPS template and submitted through OpenReview.
- All accepted contributions will be presented during the poster sessions.
- A selected number of submissions will also be invited for contributed talks.
- Submissions are non-archival and may be under review or concurrently submitted elsewhere. We especially encourage ongoing and unpublished work.
If you are excited about connecting representation learning, RL, planning, and test-time decision making, we hope you’ll join us in Sydney! 🦘
Best regards,
The Organizing Team
Ping-Chun Hsieh (NYCU), Kuang-Huei Lee (Google DeepMind), Bo Dai (Georgia Tech), Yen-Ling Kuo (UVA), Georgia Chalvatzaki (TU Darmstadt), Karen Leung (UW / NVIDIA), Co Yong (NTU), Claas Voelcker (UT Austin)
**CALL FOR PARTICIPATION**
ICTCS 2026 - 27th Italian Conference on Theoretical Computer Science
Udine, Italy, September 7-9, 2026
The 27th Italian Conference on Theoretical Computer Science (ICTCS 2026) will take place in Udine, Italy, on September 7-9, 2026.
ICTCS is the conference of the Italian Chapter of the European Association for Theoretical Computer Science (EATCS). It provides a forum for researchers working across the broad spectrum of theoretical computer science to exchange ideas and results, and offers an opportunity for junior researchers and PhD students to interact with senior members of the community.
The list of accepted papers is now available on the conference website. The detailed conference program will be published shortly.
### Invited Speakers
We are pleased to welcome the following invited speakers:
* Parosh Aziz Abdulla - Uppsala University, Sweden
* Quantum Circuit Verification: Challenges and Opportunities
* Nadia Pisanti - University of Pisa, Italy
* Alignments to and among Degenerate Strings
### Registration
Registration for ICTCS 2026 is open. The registration deadline is August 31, 2026.
The registration fee includes access to all conference sessions, coffee breaks and lunches, the conference dinner, the social event, and one-year membership in EATCS and in the Italian Chapter of EATCS.
### Venue
The conference will be held at Sala Tomadini, Via Tomadini 30, Udine, within walking distance of the city centre.
For the list of accepted papers, registration details, travel information, and forthcoming updates on the program, please visit the ICTCS 2026 conference website:
https://ictcs2026.uniud.it/
We look forward to welcoming you to Udine in September!
The ICTCS 2026 Organizing Committee
Inviato da Outlook per Android<https://aka.ms/AAb9ysg>
==> Only **one more week** to submit your abstract! <==
*** FINAL Call for Abstracts ***
*** NARNiHS 2027
*** North American Research Network in Historical Sociolinguistics
*** Ninth Annual Meeting
*** 100% IN PERSON
*** Co-Located with the Linguistic Society of America (LSA) Annual Meeting
*** San Francisco, California, USA
*** 6-9 January 2027
This event offers an opportunity for scholars from all over the world to gather and share leading research in historical sociolinguistic methods of linguistic inquiry; and we invite our fellow historical sociolinguists and scholars in related fields from our global scholarly community to join us in San Francisco for our 9th Annual Meeting.
Consult this Call for Abstracts on the web: https://narnihs.org/?page_id=3542 .
---------- Call for Abstracts ----------.
Abstract submission online:
https://easyabs.linguistlist.org/conference/NARNiHS_2027
Deadline: Monday, 24 August 2026, 11:59 PM US Eastern time.
Late abstracts will not be considered.
The North American Research Network in Historical Sociolinguistics (NARNiHS) is accepting abstracts for its 9th Annual Meeting. As a Sister Society of the Linguistic Society of America (LSA), NARNiHS will hold its Annual Meeting sessions at the LSA Annual Meeting in San Francisco, Wednesday-Saturday, 6-9 January 2027. The 9th edition of this inclusive NARNiHS event seeks to provide a collaborative environment where presenters bring fully developed work for presentation and enrichment. We see the NARNiHS Annual Meeting as a place for showcasing excellent projects in historical sociolinguistics, seeking feedback from peers, and engaging in productive development of the field's enduring questions.
NARNiHS welcomes papers in all areas of historical sociolinguistics, which is understood as the application and/or development of sociolinguistic theories, methods, and models for the study of historical language variation and change over time, or more broadly, the study of the interaction of language and society in historical periods and from historical perspectives. Thus, a wide range of linguistic areas, subdisciplines, methodologies, and adjacent disciplines easily find their place within historical sociolinguistics, and we encourage submission of abstracts that reflect this broad scope.
Abstracts will be accepted for both 20-minute papers and posters. Accepted papers will be organized into thematic sessions that will be scheduled throughout the 4 days of the LSA conference; accepted posters will be scheduled together in the LSA poster gallery. Please note that, at NARNiHS events, poster presentations are an integral and equal part of the intellectual exchange (not second-tier presentations). Abstracts will be assigned a paper or a poster presentation based on determinations in the review process about the most effective format for the submission. However, *if you prefer that your submission be considered primarily for poster presentation*, please specify this in your abstract.
Successful abstracts will demonstrate *thorough grounding* in historical sociolinguistics, *scientific rigor* in the formulation of research questions, and *promise for rich discussion* of ideas. Successful abstracts will be explicit about which *theoretical frameworks*, *methodological protocols*, and *analytical strategies* are being applied or critiqued. *Data sources and examples* should be sufficiently presented, so as to allow reviewers a full understanding of the scope and claims of the research. Please note that *the connection of your research to the field of historical sociolinguistics* should be explicitly outlined in your abstract. Failure to adhere to these criteria will likely result in rejection of the abstract.
*** Abstract Format Guidelines ***.
- Abstracts must be submitted in PDF format.
- Abstracts must fit on one 8.5x11-inch or A4 page, with margins no smaller than 1 inch (2.54 cm) and a font style and size no smaller than Times New Roman 12 point. You are encouraged to use the entire page to provide a full and robust description of the research. All additional supporting content (visualizations, trees, tables, figures, captions, examples, and references) must fit on a single (1) additional page. No exceptions to these requirements are allowed; abstracts longer than one page or with more than one additional page of supporting content will be rejected without review.
- Specify if you prefer your submission be considered primarily for a poster presentation.
- Anonymize your abstract. We realize that sometimes complete anonymity is not attainable, but there is a difference between the nature of the research creating an inability to anonymize and careless non-anonymizing (in citations, references, file names, etc.). Be sure to also anonymize your PDF file (you may do so in Adobe Acrobat Reader by clicking on "File", then "Properties", removing your name if it appears in the "Author" line of the "Description" tab, and re-saving the file before submission). Do not use your name as the file name when saving your PDF (e.g. Smith_Abstract.pdf); file names may not be automatically anonymized by the EasyAbs system. Rather, use non-identifying information in your file name (e.g. HistSoc4Lyfe.pdf). Your name should only appear in the online form accompanying your abstract submission. Abstracts that are not sufficiently anonymized wherever possible will be rejected without review.
*** General Requirements ***.
- Abstracts must be submitted electronically using the following link: https://easyabs.linguistlist.org/conference/NARNiHS_2027
- Authors may submit a maximum of two abstracts: One single-author abstract and one co-authored abstract.
- Authors may not submit identical abstracts for presentation in NARNiHS Sessions and at the LSA Annual Meeting or another LSA Sister Society meeting (ADS, ANS, NAHoLS, SCiL, SPCL, or SSILA).
- Specify if you prefer your submission be considered primarily for a poster presentation.
- After submission, no changes of author, title, or wording of the abstract may occur. If your abstract is accepted, adjustment of typographical errors is permitted before a final version of the abstract is printed in the conference booklet.
- Papers and posters must be delivered as projected in the abstract or represent bona fide developments of the same research.
- Authors are expected to attend the conference in-person and present their own papers and posters. This will not be a hybrid event.
Contact us at NARNiHistSoc(a)gmail.com with any questions.
Call for papers
Special Session on Agentic AI for Spatio-Temporal Information Extraction
and Reasoning - SAIS 2026
Following the rapid emergence of agentic applications and the growing
volume of event-centric news and social media posts, we announce the
first special track on Spatio-Temporal Information Extraction and
Reasoning of the AGENTICS 2026 conference, which this year is part of
the 18th Joint Conference on Computational Intelligence and will take
place in Angers, France.
We aim to bring together research where agentic AI and Large Language
Models are utilized for spatio-temporal grounding of events, entities,
and sentiments from textual sources. This also includes the construction
of event and entity relations, graphs, and timelines and the
spatio-temporal inference of implicitly mentioned locations and time
periods.
We welcome contributions on agentic reasoning for spatio-temporal task
planning: for example, spatio-temporal reasoning from information
extracted from text; reasoning over structured event databases,
spatio-temporal knowledge graphs, and clusters; spatio-temporal
attribute propagation; reasoning about sequences of events; and all
other spatio-temporal reasoning in a text-based context.
We will address also downstream applications of these technologies like
domain-specific monitoring agents for conflict, disasters, environmental
change, public health, and other crisis and policy-related domains.
Finally, we encourage work on early warning and forecasting with LLMs,
and on reasoning over event sequences and spatial proximity to support
downstream analysis and decision-making.
In particular, we welcome submissions addressing the following main
topics: spatio-temporal grounding and extraction using or interacting
with LLM or agentic AI, spatio-temporal agentic reasoning, as well as
downstream agentic or LLM-based applications. We also address evaluation
techniques and benchmarks of the above mentioned approaches.
Additionally, submissions on all other topics, related to agentic and
LLM based spatio-temporal reasoning and extraction from texts will be
considered.
Website: https://agentics.scitevents.org/SpecialSessions.aspx
Important dates:
Paper Submission: September 1, 2026
Authors Notification: September 9, 2026
Camera Ready and Registration: September 18, 2026
Co-chairs of the Special Session:
Hristo Tanev
Joint Research Centre, Data Intelligence for Policy Unit, European
Commission, Italy
Milena Slavcheva
Institute of Information and Communication Technologies, Department of
Artificial Intelligence and Language Technologies, Bulgarian Academy of
Sciences, Bulgaria
Bertrand De Longueville
Joint Research Centre, Text Mining and Artificial Intelligence
Competence Centre, European Commission, Italy
Call for participation → Shared Task: ASR and Language Resources for Indigenous Languages of South America (SIMBig 2026, Peru) | https://simbig.org/SIMBig2026/en/sharedtask.html
Deadline: 21/09/2026 to participate in the shared task (👉 Register: https://forms.gle/112jyHYPq9SVMmvs5)
We invite submissions for the SIMBig 2026 Shared Task on ASR and Language Resources for Indigenous Languages of South America, a unique opportunity to contribute to low resource language preservation while competing for publication in Springer CCIS.
South America is home to a breathtaking linguistic tapestry, where hundreds of Indigenous languages are spoken, each carrying unique worldviews, ancestral knowledge, and cultural identities. Yet, this rich heritage is at a risk of being left behind. As the world becomes increasingly digital, the survival of these languages faces a new challenge: a deep technological gap.
Why participate? We have prepared three tasks to accommodate different interests:
Task 1: Puno Quechua ASR baseline (we provide a 66-hour speech corpus)
Task 2: ASR for any Indigenous South American language without an existing ASR
Task 3: Educational/literacy resource based on language datasets (websites, tutorials, notebooks, etc.) for any Indigenous South American language(s).
Publication opportunity: Winning teams receive cash prizes, free registration to SIMBig 2026, and per-review publication in the Springer Communications in Computer and Information Science (CCIS) series.
Important dates (all AoE):
24/08/2026 - Release of datasets (Scripted Speech, Spontaneous Speech).
16/09/2026 - Open orientation hour (Register: https://forms.gle/FnL2Xp8X5j3zCsHZ9)
21/09/2026 - Registration deadline
30/09/2026 - Release of test dataset for Task 1
18/10/2026 - Submission deadline for final results and system description paper
28-30/10/2026 - Results announcement during SIMBig 2026
For any inquiries, please contact the organiser, Elwin Huaman (elh97(a)cam.ac.uk)
Dear All,
At the turn of 2026 and 2027, the next edition of PolEval, a shared task competition for computational tools for processing Polish, will take place. The aim of the shared task is to improve the quality of existing solutions, test new algorithms and methods, and support the development of computational linguistics.
We are now opening the Call for task proposals for PolEval 2026. Until September 7, 2026, we invite submissions from individuals, teams, and organisations interested in preparing and coordinating shared tasks as part of the competition.
A task proposal should include:
task title – a short and unambiguous name;
task description – outlining its goal, scope, and main assumptions;
motivation and relevance – in particular, information on whether a similar task has already been organised as part of other competitions or benchmarks, e.g. for other languages; if the proposal concerns a new problem or a problem specific to Polish, please highlight its uniqueness and potential importance for the development of natural language processing;
task preparation status – in particular, information on the availability and status of training and test data, any licensing restrictions, and the expected date by which the complete dataset will be ready;
evaluation procedure – a description of how systems will be evaluated, including the proposed metrics, result submission procedure, etc.;
proposed timeline for running the task as part of PolEval 2026 – in particular, the expected dates for data release and completion of the evaluation.
Task organisers will be responsible for preparing the task description, data, and evaluation procedure; communicating with participants; coordinating the review process for submitted papers; and publishing a paper summarising the task in the PolEval proceedings in the ACL Anthology <https://aclanthology.org/venues/poleval/>.
Task proposals should be sent to: lukasz.kobylinski(a)ipipan.waw.pl <mailto:lukasz.kobylinski@ipipan.waw.pl>
PolEval 2026 schedule
Call for task proposals: August 11, 2026
Task submission deadline: September 7, 2026
Notification of task acceptance: September 14, 2026
First call for participation/papers & release of Train and Test A data: October 13, 2026
Second call for participation/papers: November 13, 2026
Release of Test B data: December 1, 2026
Announcement of final results (Test B): December 8, 2026
Paper submission deadline: December 15, 2026
Notification of paper acceptance: January 5, 2027
Camera-ready paper due: January 19, 2027
Presentation of the results and publication of the PolEval proceedings in the ACL Anthology: March 2027
The following tasks have been organised in previous editions of PolEval:
Part of Speech Tagging
Sentiment Analysis
Dependency Parsing
Named Entity Recognition
Language Models
Recognition and normalization of temporal expressions
Lemmatization of proper names and multi-word phrases
Entity linking
Machine translation
Automatic speech recognition
Automatic cyberbullying detection
Post-editing and rescoring of automatic speech recognition results
Morphosyntactic tagging of Middle, New and Modern Polish
Word sense disambiguation
Information extraction and entity typing from long documents with complex layouts
Punctuation restoration from read text
Evaluation of translation quality assessment metrics
Post-correction of OCR results
Question answering challenge
Punctuation prediction from conversational language
Abbreviation disambiguation
Passage retrieval
Reading Comprehension
Emotion and sentiment recognition
Polish Automatic Speech Recognition Challenge
Spotting Machine-Generated Text from Language Models for Polish (ŚMIGIEL)
Gender-inclusive LLMs for Polish
Polish Speech Emotion Recognition Challenge
Across these tasks, more than 140 systems have competed, developed by researchers from academia and by private companies. Organising a task is a great opportunity not only to test and compare existing solutions, but also to draw the attention of the research and engineering communities to important challenges in Polish language processing.
More information about the competition is available on:
the PolEval website <http://poleval.pl/>;
the PolEval Discord server <https://discord.gg/dpn94tUSyT>, used for ongoing discussions and information exchange;
PolEval on LinkedIn <https://www.linkedin.com/company/18152015>.
We would also greatly appreciate your help in spreading the word about PolEval 2026 among your colleagues, students, PhD students, and others interested in computational linguistics and NLP.
Thank you in advance for your support and involvement. We look forward to receiving interesting proposals for tasks for PolEval 2026!
Best regards,
Łukasz Kobyliński
Maciej Ogrodniczuk
Alina Wróblewska
Institute of Computer Science, Polish Academy of Sciences
We invite submissions to the topical collection "Participatory AI: Co-Designing Sociotechnical Systems", published in the Springer Nature journal AI and Ethics.
(New) Submission deadline: October 7, 2026
Journal CFP: https://link.springer.com/collections/fehgbcihbb
This collection explores how participatory approaches can address risks and limitations of AI-powered technologies by engaging diverse stakeholders in the design process. Drawing on the sociotechnical tradition, it brings together work at the intersection of sociotechnical studies and participatory design, investigating how AI and digital systems can be co-designed to reflect shared values, accountability, and agency.
Topics of interest include (but are not limited to):
- Methods, frameworks, and design solutions for participatory AI (co)design
- Experiments, simulations, prototypes, or case studies of co-design in AI development
- Strategies for balancing individual and collective needs in AI design
- Critical reflections on challenges and limitations of participatory approaches
- Analyses of power dynamics and ethical considerations in participatory AI
- Experiences and lessons learned from co-design and stakeholder engagement
- Assessing AI impacts through participatory and stakeholder engagement approaches
We welcome technical and non-technical submissions with theoretical, methodological, or experimental contributions. Interdisciplinary work is explicitly encouraged. Submitted manuscripts must contain at least 80% original content.
This collection is connected to the workshop Mind the AI GAP: Co-designing sociotechnical systems (HHAI 2025), but is open to all contributions aligned with the scope.
Guest Editors:
Costanza Alfieri, University of L'Aquila
Marta Marchiori Manerba, University of Turin
Rui Prada, Instituto Superior Técnico, Universidade de Lisboa
Beatrice Savoldi, Fondazione Bruno Kessler
Ilaria Tiddi, Vrije Universiteit Amsterdam
As part of the celebration of the launch of the new LANA Corpus of spoken conversational American English, a series of free to attend events are being held at Hong Kong Polytechnic University on the 16th and 17th September 2026. On the 16th, an event in the morning will launch the LANA corpus, present details of how to download it, discuss its compilation and highlight some early results from the corpus. People can register to attend this event in person or online.
In the afternoon of the 16th, Prof. Vaclav Brezina will provide a workshop on the use of the #LancsBox tool - this is an in person only event.
On the 17th, a full day of talks on corpus-based research will take place with registration available for both in person and online attendance.
To register for the event, visit: https://polyu.hk/RstBb<https://t.edm.polyu.edu.hk/activities_web/track/click?msgid=fae75f5b-88dd-4…>
Face to face attendees will find details of the venue (Hung Hom Bay Campus) on the registration page. Online attendees will be sent joining instructions after registration. Details of the events are given below. Note that the timings given are for Hong Kong. We hope to see colleagues old and new both in person and online next month,
Tony McEnery
16th September - morning, LANA Launch Event
9:00-9:15 Introduction, Prof. Anthony McEnery, The Hong Kong Polytechnic University
9:15-10:15 Building LANA-CASE: Spoken Corpus Design, Compilation, and Lessons Learned, Dr Elizabeth Hanks
10:15-11:15 Is Casual Conversation Linguistically similar to Other Spoken Registers? Prof. Tove Larsson, Northern Arizona University
11:15-11:45 Coffee Break
11:45-12:45 and I was like wait what? The Language of Gen Z Americans, Prof. Paul Baker, The Hong Kong Polytechnic University
12:45-13:00 Accessing LANA, Prof. Anthony McEnery, The Hong Kong Polytechnic University & Prof. Vaclav Brezina, Lancaster University
16th September - afternoon 14:00-17:00
#LancsBox Workshop, Prof. Vaclav Brezina, Lancaster University
17th September - all day, symposium
9:00-10:00 Evaluating Chatbot Authenticity in Simulations of Spoken Interaction, Prof. Vaclav Brezina, Lancaster University
10:00-11:00 Changes in the Register of Conversation over Time - a Discourse Unit Approach, Prof. Anthony McEnery, The Hong Kong Polytechnic University
11:00-11:15 Coffee Break
11:15-12:15 Using Corpora to Analyse Health Communication: A Study of the R/Dementia
Subreddit, Dr Gavin Brookes, Lancaster University
12:15-13:15 When Does a Clause Become a Story? Tracing Discourse Across Languages and Centuries, Dr Hanna Schmück, University of Augsburg
13:00-14:00 Lunch Break
14:00-15:00 Surprise Words and Core Words: Using Vibe Coding to Identify Collocational Divergence Across Corpora, Prof. Paul Baker, The Hong Kong Polytechnic University
15:00-15:15 Coffee Break
15:15-16:15 Comparing and Contrasting in Corpus Linguistics, Dr Niall Curry, University of Birmingham
16:15-17:15 Getting on the Same Wavelength: Synergising Text, Sound, and Vision in Corpus Design, Dr Samuel Schmück, Lancaster University