Analyzing Code-Switching Patterns in Multilingual Urban Broadcast News Across Regions Using Speech-Driven Language Identification
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the study
- 1.3Problem Statement
- 1.4Objectives of the study
- 1.5Limitations of the study
- 1.6Scope of the study
- 1.7Significance of the study
- 1.8Structure of the research
- 1.9Definition of terms
Chapter TWO
LITERATURE REVIEW
- -
- 2.1Theoretical framework
-
- 2.2Review of code-switching literature: linguistic theories and sociolinguistic perspectives
-
- 2.3Methods and models for language identification in multilingual contexts
-
- 2.4Speech signal properties and phonetics relevant to code-switching detection
-
- 2.5Code-switching in media and broadcast journalism
-
- 2.6Language contact phenomena across regions and communities
-
- 2.7Sociolinguistic factors influencing code-switching (age, gender, education, urbanization)
-
- 2.8Multimodal cues in broadcast discourse (prosody, intonation, rhythm)
-
- 2.9Corpus creation and annotation for multilingual broadcasts
-
- 2.10Ethical considerations and data privacy in language research
Chapter THREE
RESEARCH METHODOLOGY
- -
- 3.1Research design and paradigm
-
- 3.2Source material and data collection plan
-
- 3.3Corpus construction: transcription, alignment, and annotation schema
-
- 3.4Language identification methodology: acoustic features and classifier choice
-
- 3.5Code-switching annotation scheme (types and triggers)
-
- 3.6Data preprocessing and quality assurance
-
- 3.7Experimental setup and parameter tuning
-
- 3.8Validation and reliability checks
-
- 3.9Ethical approval and consent procedures
-
- 3.10Limitations and potential biases
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- -
- 4.1Descriptive statistics of the dataset
-
- 4.2Quantitative analysis of code-switching frequency and distribution
-
- 4.3Temporal patterns of code-switching across broadcast segments
-
- 4.4Region-wise comparison of code-switching norms
-
- 4.5Prosodic and phonetic correlates of code-switching
-
- 4.6Language identification performance and error analysis
-
- 4.7Sociolinguistic correlates: speaker profiles and audience reception indicators
-
- 4.8Qualitative discourse analysis of selected broadcasts
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- -
- 5.1Summary of findings
-
- 5.2Theoretical contributions to linguistics and communication studies
-
- 5.3Implications for multilingual media practices
-
- 5.4Practical applications: automated moderation, subtitling, and language policy support
-
- 5.5Limitations and recommendations for future research
-
- 5.6Conclusion and final reflections
Project Abstract
This study investigates code-switching patterns in multilingual urban broadcast news across regions by leveraging speech-driven language identification to map language use, switch points, and linguistic contexts with high spatial granularity. Drawing on a cross-regional corpus of broadcast news segments from five linguistically diverse urban centers, the research integrates automated audio processing, acoustic-phonetic features, and state-of-the-art neural language identification models to detect language boundaries in real time, including rapidly switching segments, intra-sentential shifts, and lexical borrowings. The methodology combines speaker diarization, robust speech separation, and language-aware automatic transcription to construct a dense, time-stamped multilingual transcript annotated with language tags, code-switch types (within-sentences, inter-sentences, and tag-switching), and sociolinguistic variables such as region, topic, reporter background, and audience demographics. We introduce novel feature representations that incorporate prosodic cues (intonation, cadence, stress), lexical cohesion patterns, and channel-specific broadcasting conventions to improve the sensitivity and precision of language identification in noisy broadcast environments. The study employs a multi-task learning framework that jointly models language identification, speaker attribution, and switch-point detection, enabling better disambiguation when languages share lexical or phonological similarities. Through systematic analysis, we examine regional variation in code-switching frequency, preferred language pairings, and domain-specific usage (politics, economics, social issues), as well as temporal dynamics across broadcasting schedules and seasonal events. The research also analyzes the functional roles of code-switching in news discourse, such as stance marking, audience targeting, and topic emphasis, and assesses whether switching patterns correlate with audience reach, segment length, or perceived credibility. An evaluative component compares automated outputs with manually annotated gold standards to establish reliability metrics across languages with varying resource availability and dialectal diversity. The study contributes to methodological advances in multimodal language monitoring by providing a reproducible pipeline that integrates acoustic signals, textual transcriptions, and sociolinguistic metadata, along with an annotated corpus of multilingual broadcast content suitable for benchmarking speech-driven language identification in real-world media settings. Practical implications are discussed for media production, audience analytics, and editorial decision-making, particularly in multilingual societies where timely, accurate language-aware reporting can enhance accessibility and inclusivity. Limitations include potential biases arising from unequal language representation, dialectal variation, and the quality of archival audio. The anticipated outcomes include a scalable framework for fine-grained code-switching analysis, a publicly available multilingual broadcast dataset, and insights into regional linguistic dynamics that influence information diffusion and public perception in the media landscape.
Project Overview
What This Project Is About
A straightforward exploration of how multilingual speakers switch between languages in urban TV news and how technology can detect those switches by listening to the spoken content.
The Problem It Addresses
Many urban newsrooms broadcast in more than one language, but we lack easy methods to track when and why speakers switch languages. This project fills that gap by studying patterns of code-switching to support better understanding of multilingual audiences and improve language-aware news analytics.
Objectives of the Project
- Describe common code-switching patterns in multilingual broadcast news.
- Introduce a simple method to detect language switches from spoken news audio.
- Compare switching behavior across regions or cities.
- Assess how switch points relate to segment topics or audience practice.
What You Will Do Step by Step
1. Gather publicly available broadcast news clips from multiple regions. 2. Transcribe a subset of the audio into text. 3. Label language segments where possible. 4. Apply a basic language-identification tool to detect switches. 5. Analyze how switch points align with topics or regions. 6. Summarize findings and discuss implications for journalism and language studies.
Expected Outcome
We expect a clear outline of typical code-switching patterns, a simple workflow for detecting switches in speech, and initial insights into regional differences that could guide multilingual newsroom practices and future research.