Skip to main content

Chao, Amber_Senior Thesis 2026

Page 1


A Preliminary Analysis of the Brain Treebank Dataset

Amber Chao

Senior Thesis | 2026

AmberA.Chao

BostonUniversityAcademy,BostonUniversity

SeniorThesis

SrdjanDivac

January20,2026

Abstract

Thispaperpresentsanoverviewofandresearchutilizing theBrainTreebankdataset.BrainTreebankisadatasetof neurological,electroencephalographic(EEG)recordingstaken frommovie-watchingtestsubjects,providedbytheMITComputer ScienceandAILab.Thedatasethasasamplesizeof10subjects,a cumulativerecordingtimeof43hoursacross26movies,andis currentlythelargestdatasetofintracranialresolutionwith linguisticannotations(Wangetal.,2024).Thus,researchwas conductedtoassesstheusabilityofthenovelBrainTreebank datasetbyutilizingthedatasettoconductlanguagedecoding research.Naturallanguagedecodingisaprocessthatisimportant fordevelopingbrain-computerinterfacesandessentiallyconsists ofutilizingcomputationalmethodstodecipherspokenorheard wordsandsentencesfromneurologicalrecordings.Theresearch presentedinthispapersoughttodecodeheardlanguagefrom neurologicalrecordingsintheBrainTreebankdataset.Usingthe Pythonprogramminglanguage,thedatasetwasanalyzedforthe presenceofpreviouslyknownlinguisticneurologicalfeatures,such astheN400.Whiletheresultswerenotsignificantenoughto classifytheproducedmodelsasaccuratenaturallanguage decoders,themodelswereabletomorereliablyidentifybasic syntaxandsentencestructure(Wangetal.,2024).Therefore,these resultsdemonstrateBrainTreebank’susabilityforneurolinguistic researchandimpliesarequirementforamoreprecisedataset,a largerdataset,orbettermethodsofanalysistocreateatrue languagedecodingmachine.

Introduction

TheInfoLabattheMITComputerScienceandAILab (CSAIL)isonthecuttingedgeoflanguageresearchtopushthe boundariesofhuman-machinecommunication.Thenecessityof effectivecommunicationbetweenhumansandcomputershasonly grownasartificialprocessingcapabilityincreased,andthenext generationofsupercomputer-boundlargelanguagemodelsis alreadychallengingourbandwidthofcommunication(Jinetal., 2025;Tamkinetal.,2021).Inotherwords,ourabilityto comprehendandprocessthedataoflargelanguagemodelsisfar outstrippedbytheirsheermagnitude.Apossiblesolutioncould involvefacilitatingdirectcommunicationbetweenhumanand artificialbrainswithabrain-computerinterface.Theresearchin thispaperhasapplicationsinlanguage-basedbrain-computer interfaces,andispoisedtopushforwardscientificunderstanding ofthebrainaswellasitsdigitalcounterparts.However,thecore focusofthisresearchistobetterunderstandhumanlanguage.

Brain:OriginsandAdvancementsinNeurologyandBrain Mapping

Aslanguageisoneofthemostcomplexanduniquehuman adaptations,researchaimedatunderstandingithasnaturally focusedonthebrainovertime.Modelsofthehumanbrainhave evolvedasscientificunderstandinghasincreased,beginningwith basicmonolithicmodelswhichtreatedthebrainasasortof black-boxandexpandingtomorenuancedmodelsthatrecognized thebrainasbeingmadeupof(mostly)discretestructureswith specificpurposes(Hartwigsen,2024).Forinstance,inthelate19th century,PaulBrocaandCarlWernickediscoveredareas responsibleforlanguageproductionandlanguagecomprehension, respectively,bylocalizingareasofthebrainaffectedby stroke-inducedlesionsinpatientswithpatientswitheachsymptom (Hartwigsen,2024).Thus,researchintothelinguisticsignificance ofspecificbrainareaswasinitializedbyBroca,Wernickie,and

eventuallyGeschwind,andeachoftheirdiscoveredbrainareasare noweponymouslynamed(Hartwigsen,2024).

However,scientificunderstandingofthebrain’sstructures andfunctionsgreatlyincreasedwiththeadventofmorecomplex methodsofneuroimaging,suchasmagneticresonanceimaging (MRI),functionalmagneticresonanceimaging(fMRI),and magnetoencephalography(MEG).Magneticresonanceimagingis atechnologythatchieflymapsthedistributionandbondedstateof hydrogenatomsthroughoutavolume(Groveretal.,2015).Since areasofwhitematterandgreymatterinthebraindifferinfatand waterconcentrations,theyprovideMRIimagesofexcellent contrast(Langkammeretal.,2012).Giventheseproperties,MRI hasenabledscientiststogeneratethree-dimensionalmapsofbrains invivo,apossibilityunthinkabletoBrocaandWernicke. FunctionalMRIshrinksthesampletimeofconventionalMRIto withinanorderofmagnitudeofonesecond,allowingreal-time monitoringofbrainfunctions(Huotarietal.,2019).Thus,fMRIis mostusefulfortrackingtheflowofbloodandcompoundswithin thebloodthroughthebrain,suchashemoglobin(Huotarietal., 2019).AsMRIcandistinguishoxygenatedanddeoxygenated hemoglobin,fMRIcandetectwhichareasofthebrainareactive duringdifferentactivitiesbasedontheirconsumptionofoxygen (Huotarietal.,2019).Magnetoencephalographyisaprocesswhich directlymeasuresthestrengthofmagneticfieldsthroughthe volumeofthebraintopinpointneuronalactivity.Assuch,MEGis wellsuitedtolocalizingsourcesofbrainactivity,especiallythose causedbyseizures(Baillet,2017).WhileMEGisthemostcapable technologyofthethreeforstudyingcognitiveprocesses,allthree fallshortwhencomparedtoelectroencephalography(EEG) (Arjoonsinghetal.,2024).

Electroencephalographyisamethodtodirectlymeasure theelectricalimpulsesgeneratedbytheactivityofneuronswithin thebrain.Thisisincrediblyusefulfortheresearchbeing

conductedwithinthispaperbecauselinguisticconstructssuchas semantics,phonology,andsyntaxaremostpresentwithinthebrain asintracranialfieldpotentials(IFPs)(Hartwigsen,2024).1 Foran exampleofsemanticsignificanceinneurologicalfunctions, scientistsusedextracranialelectroencephalographytodiscoverthe N400,anegativevoltagespikethatoftenoccursinthebrain400ms afterthepresentationofanunexpectedorsemantically-violatory word(Hartwigsen,2024).Whilemorebasicandcommonly-used methodsofextracranialscalpEEGareusefulformeasuring generalpatternsinbrainfunctionwithhightemporalresolution, theylackthespatialresolutiontoisolateandlocatespecificsignals throughoutthebrain(Parvizi&Kastner,2018).Thus,thenext evolutionisstereo-electroencephalography(sEEG);sEEGutilizes intracranialelectrodes,whichcanachievesignal-to-noiseratiosup to100timesbetterthanscalpEEGbybypassingconduction throughtheepidermisandskeletonaswellasahigherdegreeof localizationduetosmallerandmoreevenlyspreadelectrode patches(Parvizi&Kastner,2018).Thankstothistwo-foldbenefit, brainsignalsofhigherqualitycannowbeacquired(Parvizi& Kastner,2018).UsingthebenefitsofsEEG,thisstudyaimsto furthertheneurologicalunderstandingoflanguageprocessing.

Treebank:AToolofLinguisticResearch

Oneofthemaintoolsofcomputationallinguisticsisa treebank,whichisadatastructurebestsuitedtothestorageand analysisofsemanticandsyntacticdata.TheUniversal Dependenciesprojectmaintainsoneofthelargest,most multilingualsetsofsyntactictreebanks,afactwhichlendsgreat credibilitytotheirprotocolsforannotation(Universal Dependencies,n.d.).Thus,tocreateadataformatsuitablefor neurolinguisticresearch,theresearchersatMITcreatedBrain Treebank:acollectionofneurologicalrecordingsand

1 Thevoltagescreatedbyneuronactivityinthebrain.

accompanyingEnglishwords,annotatedintheUniversal Dependenciesstyle.

Researchersbehindthelabforwhichthisstudywas conductedcollectedsEEGrecordingsfrommovie-watchingtest subjects.Thecollectedneurologicaldatawasannotatedwith myriadfeaturesfromthemoviethatpertainedtonaturalspeech processingsuchassentenceonsets,wordsurprisal,andpartof speech.Specifically,theannotationsfollowtheguidelinesofthe UniversalDependenciesprojectfortreebankstomaximizethe usabilityoftheneurologicaldataforlinguisticresearch.Thus,the datawasthenpackagedintoBrainTreebank,thehighestspatial resolution,linguistically-annotatedneurologicaldataset.Thisstudy aimstovalidatetheBrainTreebankdatabaseforitsaccuracyand applicabilitytonaturalspeechprocessingaswellasto preliminarilyanalyzethedatabasefornaturallanguagedecoding.

Methods

Thispaper’sresearchwasconductedutilizingdata exclusivelyfromtheBrainTreebankdatabase.TheBrainTreebank datasetwascollectedbytheInfoLabgroupattheMITComputer ScienceandAILabinconjunctionwiththeMITCenterforBrains, Minds,andMachines.BrainTreebankis,atthetimeofwriting,the largestlinguistically-annotated,stereo-electroencephalographic database(Wangetal.,2024).Thedatasetwaspublishedin2024 andwascollatedoveranundeterminedperiodoftimeprior(Wang etal.,2024).Thestudyfromwhichthedatasetwasgatheredwas conductedattheBostonChildren’sHospital(BCH)andproceeded withtheapprovalofBCHandtheHarvardIRB(Wangetal., 2024).Additionally,theclinicalresearchdataofthestudywas collectedwiththeinformedconsentofeachparticipant(Wanget al.,2024).

The10participantsofthestudyrangedfrom4to19years ofageandalreadyhadsEEGelectrodesimplantedforseizure localization(Wangetal.,2024).Eachparticipantviewed

full-length,Hollywoodmoviesona39.1centimeter,2880by1800 pixellaptopscreenplaced60to100centimetersawayfromthe participant(Wangetal.,2024).Movieswereplayedusinga customMatLabvideoplayertoensureconstantframerateplayback andaccuratesynchronizationsignalswiththeneuralrecorder (Wangetal.,2024).Additionally,movieswerere-renderedat 23.976fpstoprovidethesaidconstantframeratesourcematerial (Wangetal.,2024).Participantswerefreelyallowedtochoose moviesattheirowndiscretion,pause,resume,adjustvolume,and changepositions(Wangetal.,2024).However,participantswere askedtofocusontheplayedmoviestomaximizetheamountof usefulnatural-language-processingbraindata(Wangetal.,2024). Forthesamepurpose,participantsdidnotspeakduringthetrial andwerenotexposedtoalternate,externalsamplesofspeech whilethemoviewasplaying(Wangetal.,2024).Theexperiment proctorpausedandresumedthemovieasneededtoensurethe participant’sfullattentionwasgiventotheentiretyofeachmovie (Wangetal.,2024).

Theintracranialelectrodesutilizedinthisstudywere between21-56mmlongand0.8mmindiameter(Wangetal., 2024).Eachelectrodehadbetween6and16distinctcontact electrodesmeasuring2mminlengtheachwith1.5mmspacing betweeneachcontact(Wangetal.,2024).Thelocationatwhich eachsEEGprobewasimplantedwasdeterminedsolelybythe clinicalconsiderationsforseizurelocalization(Wangetal.,2024). Priortorecordingneurologicaldata,allparticipantsunderwenta oneTeslaMRIscantorecordthelocationsofthemetallic electrodepatches(Wangetal.,2024).FreeSurferwasthenutilized toremoveartifactsfromtheskullandparceltheelectrodesonto differentareasoftheDesikan-Killianybrainatlas(Wangetal., 2024).Thelocalfieldpotentialsmeasuredbytheprobeswere capturedwithamulti-channelneuralrecorderat2048Hz(Wanget al.,2024).Also,triggersignalsweresentfromthemovie-playing

laptopusingaUSBtriggerboxtosendcoordinatingpulsesevery 100ms,allowingtheextremelyprecisesynchronizationofmovie andneuraldata(Wangetal.,2024).Nonfunctionalelectrodes whichproducedexcessivelynoisydataforlongperiodsoftime werealsomanuallyremovedfromthedataset(Wangetal.,2024). Theresultantneuralrecordingwasscrubbedofidentifying informationsuchasictalbrainactivityandsavedtoanHDF5file (Wangetal.,2024).

Themovieswatchedbysubjectswereprimarilyannotated withtheirlinguisticfeatures,butauditoryandvisualfeatureswere alsocomputedandrecordedassupplementarydatatoavoidfalse correlations(Wangetal.,2024).Forthelinguisticfeatures,each movie’sclosed-captionsweresavedalongsidethemovietimeand splitupintoindividualwordsthatweremanuallyadjustedtoalign eachwords’onsetandoffset(exactstartandend)inthemovie (Wangetal.,2024).Stanford’spythonnaturallanguagetoolkit, calledStanza,wasfurtherusedtoparsethetranscriptofeach movietoannotatewords’partofspeechandpositioninasentence (Wangetal.,2024).Eachwords’surprisal2 wasalsocalculated usingtheGPT-2largelanguagemodel(Wangetal.,2024).Forthe non-linguisticfeatures,500mswindowsbeforeandaftereach words’onsetwereanalysedfortheirauditoryandvisualproperties (Wangetal.,2024).Auditorypropertiesincludedtheaverage pitch,changeinpitch,averagevolume,andchangeinvolumefor eachword(Wangetal.,2024).Visualfeaturesincludedthescreen brightness,numberofonscreenfaces,globalopticalflow magnitudeandangle3,andtheregularopticalflowmagnitudeand

2 Ametricwhichencapsulateshowunexpectedagivenwordisin thecontextofthesurroundingtext.

3 Ametricwhichrepresentstheamountanddirectioninwhichthe movie’scameramovesoveracertainclip.

angle4 (Wangetal.,2024).Formorequalitativevisualfeatures, cutsinthemovieweredetectedusingthePySceneDetectlibrary andeachdistinctscenewasreferencedagainstthePlaces365 datasettoproducearoughdescriptionofeachscene’slocation (Wangetal.,2024).Additionally,thespeakerofeachwordwas manuallyannotatedinthedatasetbytrainedannotatorswith specificcaretoincludeeverynameaspeakermayhave, irrespectiveofthechronologicalprogressionofthemovie(Wang etal.,2024).Whenreviewedaftercollection,thedataappearedas listsofallthewordsspokeninamoviewiththeexacttimeat whichtheywereeachspoken,alongwiththeirlinguistic,auditory, andvisualfeatures.

ResultsandDiscussion

ThisstudyshowedthattheBrainTreebankdatabasewas accurateforlanguagedecodingasotherlinguisticfeatures,suchas theN400andsentenceonsets,whichwereeasilydetectablein brainwaves.However,conclusiveresultssurroundingthe effectivenessofdecodingspeechfromneurologicalsignalscould notbeobtained–thedatayieldedalotofmixedresults.For example,linearmodelsprovidedmediocrecorrelationsbetween partsofspeechandneurologicalactivity,therebymakingitnearly impossibletodecodethepartofspeechofawordspokenata givenmovementintimefromtheneurologicaldata,muchlessthe worditself.

Thispaper’sresearchontheBrainTreebankdatabasewas conductedentirelyinthePythonprogramminglanguage.Thedata analystwasfamiliarizedwiththedatastructureandanalysis methodologythroughtheJupyterNotebookquickstartfile.All filenameswerepredicatedwiththeanonymizedsubjectidandtrial

4 Ametricwhichrepresentstheamountanddirectioninwhich objectsonscreenmovewithrespecttotheirbackgroundovera certainclip.

number,witheachtrialcontainingthedataofonlyonemovie viewing.The~.h5filescontainneuralrecordingsandthe ~timings.csvfilecontainsthesynchronizationoffsetsbetweenthe recordingsandabsolutetime(i.e.thewallclock).The ~metadata.csvfilecontainsinformationaboutthemovie,suchas thetitle,itslength,anditsmovieID(aninternalmarkerusedto differentiatebetweenmovies).The~features.csvcontainsthe transcriptofthemoviesaswellasthefeaturessurroundingeach word.The~electrode_labels.csvfilecontainsdescriptionsofthe electrodepositionspairedwithelectrodeids.Thegeneralanalysis methodologyisasfollows:Usingtheh5pylibrary,theneuraldata wasimportedintoanarrayof n rows,where n isthenumberof recordedelectrodes.Eachrowcontainsanun-timestampedarrayof voltagesthelengthoftherecording.Thenthetriggerfilewas importedusingthePandaslibraryintoanarray,whichcontained thetypeoftrigger,timeinthemovieatwhichthetriggeroccurred, thestartandendtimeofthetriggerpulseinabsolutetime,the indexoftheneuralsamplerecordedatthatcertaintrigger,andthe differencebetweenthecurrentandnextindex.Then,the features.csvfilewasimportedthesamewayasthetriggerfile, producinganarrayof42columns.Thefirstcolumncontained everywordspokenwithinthemovie,twoothercolumnscontained thestartandendofeachwordinmovietime,andtheother39 columnscontainedallthepreviously-mentionedfeaturesofthe BrainTreebankdataset.

Afterimportingallthedata,theneuralrecordingarraywas convertedintoaNumPyarray(fromtheNumPylibrary)to improvethenumberofavailablefunctionsandmakethedatamore easilymanipulable.Afterwhich,thearrayofneuralrecordingswas thenfilteredtoremoveinterferenceusingtheSciPylibrary.A notchfiltersetto60Hzandthecommonharmonicsof60Hzwas appliedtoremoveanynoisecausedbythelowfrequency, alternatingcurrentofpowertransmission.Then,theonsetofeach

11 wordwasalignedtoasampleindexoftheneuralrecordingusing themovietimeofeachwordandsynchronizationtriggerfile.After which,anarrayofeachwordandtheneuraldatasurroundingthe occurrenceofthatwordwasgenerated.Fortheremainderofthis paper,theseelementswillbereferredtoas‘wordwindows,’asitis notatedinthequickstart.Thesewordwindowscontainthe neurologicaldatasurroundingeachwordofthemovieanddefine thechunksizeofneurologicaldataanalyzedforallfutureanalysis.

Themainneurolinguisticanalysissoughttoextract evidenceoftheN400fromtheBrainTreebankdata.TheN400isa prominentevent-relatedpotential(ERP)withastatusasakey standardorcontrolforverifyingsemanticreactivityinthebrain (Kutas&Federmeier,2011).WhiletheN400wasinitiallythought tobeapartoftheP300ERP,whichcorrespondstothecognitive processofdecisionmaking,theN400waseventuallydistinguished asfollowingthepresentationofunexpected,semantically-violatory words(Shainetal.,1980).However,furtherresearchhas concludedthattheN400isactuallyamoregeneralresponseand partoflanguageprocessing,andthatthesurprisalofthepresented wordsimplymodulatestheN400(Arbeletal.,2010).

Surprisalisakeymetricforanalyzingthebehaviorofthe N400.Whilesurprisalcanberesearchedwithphysicalstudies utilizingclozetests5,thescaleoflanguagesamplesrecordedinthe BrainTreebankaswellasthewidespreadexposuretothemovies usedinthedatasetmadehumantrialsanincrediblyunappealing option(Nair&Oh,2026).Therefore,usinglarge-languagemodels (LLMs)tocalculatesurprisalwasanobvioussolution.LLMs providethemostprobablefollowingwordgivenacertain contextualinput(Nair&Oh,2026).However,itisalsopossibleto runtheinverseoperationtocalculatehowprobableagivenwordis

5Anexperimentwherewordsareomittedfromcompletesentences andparticipantsareaskedtoprovideexpectablereplacements.

tofollowacertaincontext.But,becausetheprobabilityofany singlewordfollowinganysinglecontextisexceedinglysmall, giventheextensivecorpusoftheEnglishlanguage,andnever exceedsone(a100%probability),surprisalispresentedasthe negativelogofprobability(Nair&Oh,2026).Themoreprobable awordis,thecloseritssurprisalistozero,andthelessprobablea wordis,thecloseritssurprisalstraystowardinfinity.

However,theexactrelationshipbetweensurprisaland N400amplitudeisunknown,giventhevariabilityofthebrainand theunknownpowerofthelog.Thus,theresultsofN400analysis resultshadtobemanuallyanalyzedforvisualcorrelations.Within theN400analysis,thedataanalystmeasuredneuralresponse400 millisecondsafterwordoffsetandcomparedtheinstantaneous voltageat100mstotheaveragevoltageofeachwordwindow, excludingthespike.Bydividingthespikevoltagebytheaverage voltage,theanalystmeasuredavoltagedifferencefrombaseline. Next,thecorrelationbetweenthe100msvoltagedifferenceandthe wordsurprisalwasmeasured.Directlycomparingthemagnitudeof theN400voltagespiketothewordsurprisalisquitedifficult becausefurtherresearchisneededtounderstandtheexact correlationbetweenthetwofactors.However,itisclearthata negativevoltagespikeispresentintheplaceoftheN400after wordswithhighsurprisal.Therefore,thedataanalystcreateda3D graphwithsurprisal,neuralactivity,andthetimewindowbefore andafterwhentheN400voltagespikeshouldhaveoccurred.That producedaprettysignificantresponse,meaninganoticeablespike overanorderofmagnitudegreaterthanthesurroundingareaof data.TheeasewithwhichtheN400wasresolvedfromtheBrain Treebankdatasetisstrongevidencetowardthedataset’sviability forstudyingpreviously-knownlinguisticneurologicalfeatures,as wellasevendeterminingaroughlocationofsaidlinguistic neurologicalfeatures.Therefore,givenhoweasilyfeaturessuchas theN400aredetectablewithintheBrainTreebankdataset,these

resultshighlysuggestthatother,morecomplexlinguisticanalyses shouldbepossiblewiththesamedata.

Conclusion

Theresearchinthispaperaimedchieflyatreviewingthe BrainTreebankdataset’susabilityforneurologicallinguistic researchand,secondarily,atpursuingneurolinguisticresearch, utilizingtheBrainTreebankdatasetitself.Theresearcherreviewed thepracticesofdatacollection,formatting,andannotationwhich wereemployedintheproductionoftheBrainTreebankdatasetand foundnomajorflawswithinthemethodsandproceduresused.

Next,theresearcherfollowedtheBrainTreebankdataset’s guidelinesfordataanalysisbydownloadingcopiesofthedataset andrunningdataingestiontestsinPython.Preliminaryanalysis withthegoalofidentifyingwellestablishedneurologicallinguistic featureswasconductedandprovedtobeverysuccessful;theN400 waseasilycorrelatedwithvoltagepeaksfollowingsemantically violatorywords.Thus,withthepresenceofpreexistent neurolinguisticfeatures,thereisahighlikelihoodthatother neurolinguisticfeaturesorevenlanguagedecodingcouldbe researchedusingtheBrainTreebankdataset.

Limitations

Thelimitationspresentintheexperimentofthisresearch paperaremany,andtheresultspresentedinthispaperarefarfrom exhaustive.Someexamplesofencounteredlimitationsinclude hardwarelimitations,experimentallimitations,andenvironmental limitations.

Asfarascomputationallimitations,alldataanalysisand processingwasconductedonaneighth-generationIntelprocessor, specificallyan8-coremobileI5-8350Uwith16GBofRAM. Accesstocomputersofgreatercorecounts,moreextensive processmemory,orevengraphicscardhardwareacceleration, couldhaveopenedfurtheravenuesforresearchwithlowertime andopportunitycosts.

Additionally,electrodeplacementwasoutsideofthe controlofthestudybecausetheelectrodeswereplacedwith considerationforandseizurelocalizationonly.Theelectrode placement,inessence,wasrandom,whichcausedasignificant portionofthedatatohavelowrelevancetothegoalofthestudy.

Anotherpossiblelimitationcouldbethelowsubjectcount andnarrowsubjectdemographic.Thevolumesofinformationthat arerequiredtotrainmachinelearningmodels,especiallylarge languagemodelsfareclipsedtheamountofdatathatwasanalyzed inthisstudy. Whilenotmuchisknownabouttheparticipantsin theBrainTreebankclinicalstudy,thejuvenileandhighly-plastic natureofthebrainattheagerangesofsaidsubjectscouldproveto beaseverelimitingfactor.Furtherstudieswithmoreparticipants ofmorewidelyvaryingages,couldprovidemoregeneralizable dataforamachinelearningmodel.

Afurthersimilarlimitationisthemonolinguisticnatureof theBrainTreebankdataset.Allparticipantsspoke,understood,and werepresentedmoviesinEnglishwhichcausedtheresultantBrain Treebankdataset,andthereforethisstudy,tohavelow internationalapplicability.Furthermore,BrainTreebank’s monolinguisticnatureprecludesotherinterestingresearchonthe natureoflanguageandcognitionasmultilingualanalysiswouldbe required.

Anotherconsiderationisthevariabilityofthemediathat waspresentedtotheparticipants.Whilehighmediavariability enlargeddatasetscanhelpmachinelearningmodelsdevelopmore appropriateandgeneralresponses,italsoaddsanadditional independentvariable.Theadditionofsuchindependentvariables makesdrawingdefinitescientificconclusionsfromthedataset moredifficult.

Lastly,thereisadistinct possibilitynoisefromavarietyof sourcescouldcontinuetopersistwithintheBrainTreebank dataset.Astheintracranialfieldpotentialsrecordedinthestudy

areminisculecomparedtotheelectricalcurrentscommonlyused inallpowerelectronics,thereisapossibilitythatadditional sourcesforinterferencemaynothavebeenaccountedfor, especiallygiventherecordingenvironment’surbanlocationand thepresenceofotherhigh-poweredmedicalequipment,suchas x-raysandMRImachines.

FutureResearch

Aworthwhilegoalforfutureresearchwouldbetoexpand BrainTreebanktoencompassagreaternumberoflanguages,inthe samewaytheUniversalDependency’sTreebankdoes.Thatway, moregeneralconclusionscouldbedrawn,suchasthewayhuman languageprocessingoccursasawholeandthewaysemantic conceptsareprocessedandunderstoodinthebrain.

Futurestudiesshouldaimtoconcentrateelectrodesinthe temporallobetomaximizetheamountoflinguistically-relevant datacollected.However,withthevariablestructureandfunctionof differentindividuals’brains,itisimpossibletoknowforcertain wherethemostlinguistically-activeportionsofaperson’sbrain maybe.Therefore,agoodcourseofactionforfuturestudiescould betotakeinitialrecordingsfromallpartsofanindividual’sbrain, conductsimplelinearanalyses,ascertainareastofocuson,and re-recordneuraldatawiththesamestimulusinamore tightly-monitoredareaofthebrain.

References

Arbel,Y.,Spencer,K.M.,&Donchin,E.(2010).TheN400 andtheP300arenotallthatindependent. Psychophysiology, 48(6),861–875.

https://doi.org/10.1111/j.1469-8986.2010.01151.x

Arjoonsingh,A.,Jamal,B.C.,&Ganti,L.(2024).History andEvolutionoftheElectroencephalogram. Cureus, 16(8). https://doi.org/10.7759/cureus.66385

Baillet,S.(2017).Magnetoencephalographyforbrain electrophysiologyandimaging. Nature Neuroscience, 20(3),327–339.https://doi.org/10.1038/nn.4504

Chen,X.,Wang,R.,Khalilian-Gourtani,A.,Yu,L.,Dugan, P.,Friedman,D.,Doyle,W.,Devinsky,O.,Wang,Y.,& Flinker,A.(2024a).Aneuralspeechdecodingframework leveragingdeeplearningandspeechsynthesis. Nature Machine Intelligence, 6(4),467–480. https://doi.org/10.1038/s42256-024-00824-8.

Chen,X.,Wang,R.,Khalilian-Gourtani,A.,Yu,L.,Dugan, P.,Friedman,D.,Doyle,W.,Devinsky,O.,Wang,Y.,& Flinker,A.(2024b).Aneuralspeechdecodingframework leveragingdeeplearningandspeechsynthesis. Nature Machine Intelligence, 6(4),467–480. https://doi.org/10.1038/s42256-024-00824-8.

Grover,V.P.B.,Tognarelli,J.M.,Crossey,M.M.E.,Cox, I.J.,Taylor-Robinson,S.D.,&McPhail,M.J.W.(2015). MagneticResonanceImaging:PrinciplesandTechniques: LessonsforClinicians. Journal of Clinical and Experimental Hepatology, 5(3),246–255. https://doi.org/10.1016/j.jceh.2015.08.001

Hartwigsen,G.(2024).NeuroscienceofLanguage. Open Encyclopedia of Cognitive Science. https://doi.org/10.21428/e2759450.9793d1ae.

Huotari,N.,Raitamaa,L.,Helakari,H.,Kananen,J., Raatikainen,V.,Rasila,A.,Tuovinen,T.,Kantola,J., Borchardt,V.,Kiviniemi,V.J.,&Korhonen,V.O.(2019). SamplingRateEffectsonRestingStatefMRIMetrics.

Frontiers in Neuroscience, 13.

https://doi.org/10.3389/fnins.2019.00279

Jin,J.,Zhang,Y.,Xu,R.,&Chen,Y.(2025). An innovative brain-computer interface interaction system based on the large language model.https://arxiv.org/abs/2502.11659

Kutas,M.,&Federmeier,K.D.(2011).ThirtyYearsand Counting:FindingMeaningintheN400Componentofthe Event-RelatedBrainPotential(ERP). Annual Review of Psychology, 62(1),621–647.

https://doi.org/10.1146/annurev.psych.093008.131123

Langkammer,C.,Krebs,N.,Goessler,W.,Scheurer,E., Yen,K.,Fazekas,F.,&Ropele,S.(2012).Susceptibility inducedgray–whitematterMRIcontrastinthehuman brain. NeuroImage, 59(2),1413–1419.

https://doi.org/10.1016/j.neuroimage.2011.08.045

Nair,S.,&Oh,B.-D.(2026). Clozing the gap: Exploring why language model surprisal outperforms cloze surprisal. https://arxiv.org/abs/2601.09886

Parvizi,J.,&Kastner,S.(2018).Promisesandlimitations ofhumanintracranialelectroencephalography. Nature Neuroscience, 21(4),474–483.

https://doi.org/10.1038/s41593-018-0108-2

Shain,C.,Futrell,R.,VanSchijndel,M.,Gibson,E., Schuler,W.,&Fedorenko,E.(1980).Evidenceofsemantic processingdifficultyinnaturalisticreading. Reading Senseless Sentences: Brain Potentials Reflect Semantic Incongruity, 207(4427).

https://doi.org/10.1126/science.7350657

Tamkin,A.,Brundage,M.,Clark,J.,&Ganguli,D.(2021).

Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models.

https://www-nlp.stanford.edu/pubs/tamkin2021understandi ng.pdf

Universal Dependencies.(n.d.). Universaldependencies.org;UniversalDependencies. https://universaldependencies.org/

Wang,C.,Subramaniam,V.,Yaari,A.U.,Kreiman,G., Katz,B.,Cases,I.,&Barbu,A.(2023).BrainBERT: Self-supervisedrepresentationlearningforintracranial recordings. ArXiv (Cornell University). https://doi.org/10.48550/arxiv.2302.14367.

Wang,C.,Yaari,A.U.,Singh,A.K.,Subramaniam,V., Rosenfarb,D.,DeWitt,J.,Misra,P.,Madsen,J.R.,Stone, S.,Kreiman,G.,Katz,B.,Cases,I.,&Barbu,A.(2024). BrainTreebank:Large-scaleintracranialrecordingsfrom naturalisticlanguagestimuli. ArXiv (Cornell University). https://doi.org/10.48550/arxiv.2411.08343.

Zhang,H.,Zhou,Q.-Q.,Chen,H.,Hu,X.-Q.,Li,W.-G., Bai,Y.,Han,J.-X.,Wang,Y.,Liang,Z.-H.,Chen,D.,Cong, F.-Y.,Yan,J.-Q.,&Li,X.-L.(2023).Theappliedprinciples ofEEGanalysismethodsinneuroscienceandclinical neurology. Military Medical Research, 10(1),67. https://doi.org/10.1186/s40779-023-00502-7.

Turn static files into dynamic content formats.

Create a flipbook
Chao, Amber_Senior Thesis 2026 by Boston University Academy - Issuu