36:00 Caitlin Biggers
I love clips like this because it's not just poetry reading. It's like someone actively trying to remember on stage and like singing a little bit. And it's. I love these clips because it's a good reminder that, um, you know, this is a record of an actual event that took place. It's not just a reproduction of the poem that you could get anywhere. So that's a little aside about what I find exciting about this collection. Uh, if you can advance to the next slide, that'd be great. Zack. Um, we also built, uh, a named entity recognition, um, or ner pipeline, as Zack mentioned. Um, for those who aren't familiar, as I certainly wasn't when I started this project, NER is a large language model based technology that enables text to facilitate data analysis. Um, basically, the technology uses AI to identify people, places and things, um, and associate them with a standardized concept. For example, our NER pipeline was trained to identify concepts like emotion or nature or even more specific things like city names. And we can then use the labels to identify other recordings in our collection that also, uh, discuss these topics so you can get a sense of themes across the entire collection. That way, even if the person hadn't said that exact word. Um, it can also be used to identify specific points in a given recording. When a speaker touches on a kind of abstract concept like nature, even if they don't actually say the word nature. You might get things when someone's talking about like weather or trees or something like that. Um, so it's really a lot more robust than a keyword search. Um, this slide here shows the portal. And at right you can see the concepts, um, labeled and highlighted in different colors. This is static. Um, and I have multiple labels active right now. But if you look at the bar towards the bottom, if I was to select something like only emotion, uh, as the, as the concept that I wanted to view, the bar at the bottom would show colored dots at all the points along the specific recording. Um, when an author touches on emotion and you can do that for any of the terms that, uh, it identified. So to me, it really helps parse an audio file and kind of understand the whole thing in terms of, um, concepts, without having to listen to the entire thing and make notes on what authors talk about yourself. Um, so that's huge. Uh, between the authenticated transcript, the index and the Na data portal, we aim to kind of help people get into these recordings. Um, and I think that has been pretty successful. Um, we've, we've since this has been online for, uh, you know, more than six months now, we've had some press and then we've had really positive feedback. People love that. You can use these chapters in the different ways that you can research the collection. Um, and I've heard that it helps people kind of make connections across the collection as a whole. That would be very difficult to with, with just single videos up online, something like that. Um, so we hope that we've produced something that preserves these incredible literary voices and just generally helps people use them. Whether you're a professional researcher or someone, um, who has a casual interest in in poetry or the why or any other topic that this might touch on. Um, so I'm not the expert on GnRH. So if Zack wants to add anything on the research portal, I think you should certainly do that here. Um, but yeah, that's that's my spiel about our project.
39:36 SPEAKER_S4
Yeah, absolutely. I think the only thing I would add, you know, I think you you nailed it is, is, uh, you know, I can just show it in action. Um, uh, very, uh, very quickly, um, to kind of bring to life a little bit of what you just summarized. So, um, let me pull back up what we had before. Okay. So, um, here is the the research portal. You can access it if you go to the 92nd streetwise um, uh, main page. And if you go to the research portal here, uh, you can also just directly from their website, access the, the recordings. Um, but if we go to the research portal, maybe I want to find any place where, uh, where a particular person is mentioned by name. And maybe I'm interested in something like, you know, what was the role of women in literature? And normally we're kind of beholden to the specific words that are spoken. If we're doing search or it has to be in the the metadata that has been tagged, uh, manually. But if I don't really know what's in here, how might I just be able to use my own natural language? And so, um, if I just search here across the 850 plus recordings just that fast, I can find these thematically, uh, relevant, uh, excerpts. So here in a Q&A between the audience and the speaker. One of the audience members says, you know, one of the themes that emerges in a lot of the criticisms of your writing, just observations of your work, is the role of the male in a lot of it. And of course, you've argued in The Color Purple, and it wasn't really about men, it was about women. And then it goes on to mention this particular person. And so I can now click on this excerpt and go directly to that exact moment in the recording of the. And I can also then explore by those entities that, um, that Caitlyn was talking about. And so we can see that full list that's been extracted, um, in this recording. But maybe I want to find all the places where Toni Morrison is referenced, whether that is in this individual recording here, or maybe it's across the collection. And this is, this is what, uh, Caitlyn was talking about, you know, helping people navigate the collection of content, not just a single recording, but especially with research that's about what's similar, what's different, compare contrast. Uh, I maybe want to be able to, to see what the individual recordings in total tell us that a single recording wouldn't be able to alone. And so maybe I can go in here and find where Toni Morrison is referenced in other places, and you can do this across any of those types of entities. Or as we were showing just a moment ago, that just kind of natural language search. Um, now, one thing that I'll add is what we have have done, um, one is release a number of these tools as open source. So both the, uh, machine learning model for extracting named entities, we've, we've released, um, this has been, uh, starred at the very least, if not more so adopted by folks at, uh, Adobe, Microsoft, uh, AWS, um, uh, Yale, Duke University are all people that, um, have, uh, at the very least, uh, been paying attention to this and are starring this, uh, this platform, a number of people who are forking it and using it on their own. And we want to make these tools available to anyone, not just if you're using the TheirStory platform but but generally and the same thing. Um, this um, kind of interface for uh, being able to make your audiovisual collection searchable. Uh, we have also um, released as, uh, as open source. And so this was done just this past week. This would not have been possible without, uh, the commitment from the 92nd Street Y from melon. Um, and, uh, and everything that Michael and his team did to make this project successful, um, and, uh, and with kind of the learnings we had. And so we now have this available, I'll put the link in the chat, um, as well, where this package is up, uh, kind of the best of, uh, industry leading open source AI native databases that allow for that kind of thematic search. It also allows for things like having ChatGPT like conversations with collections. It was interesting. We had implemented that initially, but decided not to release that as a part of this portal just because we didn't want, um, AI interpreting, uh, works of creative, um, uh, individuals. Um, but all of that is essentially packaged up and, and something that now others can use, uh, really, really simply. Um, and so if anyone's interested in creating this type of kind of custom branded, uh, research portal for, for themselves, then, um, we'd be happy to, um, to have follow up conversations. I'll put a survey, um, into the chat as well. So if you want to have a follow up conversation with Caitlin, with Michael, with Vanessa, with myself on any of this, uh, we're happy to to chat. But this is all in terms of what you can build because it's an open source template. We have folks that have customized it in a number of different ways. Um, and so it's it's really, um, amazing. I'm excited about what's, what's possible. But again, this would not have been possible to get to this point without the 92nd Street Y and the the trust to. To build something like this together. Um, so, uh, I'll pause there, just on the technical sides, let me share my screen back to our, uh, to our slides and to some of our closing takeaways.