Posts Tagged ‘AI’

Flipping the Script on Audio Description: A True Crime Story

Wednesday, September 10th, 2025

Rather than ranting, I’m trying to be more creative. I’m taking a fresh, dramatic look at the state of AD. It’s an artistic exploration of how audio description went from a promising, groundbreaking career path to something threatened by AI and text-to-speech.

A dark noir aesthetic, a lonely microphone on a city street outlined like crime scene evidence, dramatic shadows, cinematic lighting."AD - A True Crime Story" in bold white title text, minimalist but striking with the Reid My Mind Radio logo in the bottom right corner,. Artwork generated by AI!

We’re taking on the true crime genre, but the case is about Audio Description itself.

Through storytelling, interviews, and a little suspense, we examine:
? The excitement following the CVAA and AD’s rise to prominence
? The cultural impact and personal connections AD created for the Blind community
? The streaming industry’s shift toward automation and TTS and the pushback from consumers
? What it might take to restore human-centered, high-quality AD

Do you write image descriptions?

As this is Flipping the Script, I know we have listeners and readers who are quite skilled in writing description.
Portraits and Portals – a powerful art showcase uplifting disabled artists and disaster survivors of color , are asking the community to share some of their creative and poetic descriptions for the artwork posted on the site.
They’ll include it among the access resources available for visitors with vision loss.
Just go to Portraits and Portals and visit the artist pages.
You’ll be able to provide your description for visual art pieces.
Who better than those who rock with Flipping the Script!!!

? Listen

Transcript

Show the transcript


AD: A True Crime Story

TR:
What’s up Family? It’s your brother, Thomas… Thomas Reid!
Host and producer of this here podcast. Sheesh!
If you’ve been here for a while, you know sometimes something about audio description gets on my last nerve and I choose to speak on it. Well, that hasn’t necessarily changed.
There’s a lot to be annoyed at.
Sub par AD is running wild out here.
Culturally competent audio description is still extremely lacking.
Streaming networks like Apple choose to use like 3 or four narrators for everything.
Ok, admittingly, I don’t watch that much on Apple, but I don’t think I’ve ever heard a narrator of color on any content.
And to all you Apple stans, please don’t send me one example in their defense.
AD providers still haven’t committed to working with Blind Narrators and QC specialists in a sustainable way.
If any network were going to show support to Blind narrators and QC Specialists, I would have put my money on Apple. This is why I don’t gamble y’all!
Last year, I decided that I’m not going to rant. Well I’m going to try not to rant. That doesn’t mean I’m going to be quiet. I want to find ways to say something but be creative.
So today, I present to you, my take on the true crime genre…
Take it away bro!
Narrator:
The following is a dramatization.
The purpose of this is not to blame, implicate or judge individuals. It is however an artistic expression of the frustration with the current state and foreseeable future of audio description.
While this story is fiction, it does include real people from real publicly held events.
In most cases their names have been withheld, but not the companies they represent.
— Reid My Mind Radio theme Music
— Suspenseful cinematic music
Lies, deceit, murder?
This is the case of AD, A True Crime Story!
— Audio: President Obama signing the CVAA
President Obama:
The Twenty First Century Communications and Video Accessibility Act will make it easier for people who are Deaf, blind or live with a visual impairment to do what many of us take for granted from navigating a TV or DVD menu, to sending an email on a smart phone.
… I’ll sign it now… (Applause)
— An uplifting pop tune.
AD:
I can’t tell you how excited I was when President Obama signed this legislation.
Even watching this now, gives me chills!
Narrator in conversation with AD:
Why was this so special to you?
Oh my goodness!
This was the chance I was waiting for since starting my career!
It might sound corny, but I always knew, this was my calling.
I was certain, this, was the opportunity I’d been waiting for.
Would you mind introducing yourself?
Sure, all my friends call me AD. My pronouns are They Them.
In two thousand ten, AD was only in their 20’s. Relatively early in their career, and they were bouncing back from some setbacks in the early two thousands.
My career really began in theater.
Then, in the 90’s, I got a chance to work with PBS.
I have to say, It really hurts me today as this administration is so set on destroying PBS and NPR.
I not only feel a kinship with these vital organizations because I worked with PBS, but I know what it feels like to have a target on your back.
I’ve been in that position for the past few years.
Not by this administration though. Well, at least, not yet!
AD continued at PBS and while they enjoyed that work, like others, they had dreams of reaching more people.
In two thousand ten, with the signing of the Twenty First Century Communications and Video Accessibility Act, or CVAA, AD’s future appeared to be bright. Advocates worked tirelessly to help pass legislation that would help AD and others not only secure the right to work, but broaden their career opportunities.
No longer would they be confined to limited hours on children’s television, they were going prime time!
I went from begging film makers to include me, to Hollywood specifically calling me! They were now seeing me, as a valuable asset able to enhance their stories!
And then, the streaming services came online. Oh my gosh! Just when I thought life couldn’t get any better.
This was around 2015.
Streaming companies were invested in AD and their talent.
Netflix Representative:
I think we first heard the calls for audio description
around 2014 when we were just launching into our own original content.
We spent most of 2014 doing a lot of the sort of back end work trying to prep ourselves to be able to provide audio description consistently. A lot of different research and creating different technical specifications to make sure that audio description would play back on all the many different devices through which people access Netflix.
We also created audio description style guides. We wanted to make sure that when we were ordering audio description that there was a consistent experience across that was high quality, that used good voice talent, good script writing, and that also met, like I said, the technical specifications so that we could ingest those files and then make sure that they could play back on all the different devices.
We also set up a lot of partnerships with different companies that create audio description to make sure we had a really healthy supply chain. Uh, knowing that we were launching full force into a huge volume of content, we knew that as we made a commitment to create audio description on all of our original content, that that was not going to be a trivial amount. We set up that supply chain and made sure we had good script authors and talent and recording studios, and also made sure that we had a really healthy pool of people who could QC audio description to make sure that we had had a set of ears on the files before they went live, to get ahead of any issues that might come out. So we worked on that for most of 2014 and then in 2015 we launched Daredevil without audio description. So that was definitely a miss, and we heard it from the community loud. Luckily, we had already done all that legwork, like I said, in 2014 so we recovered quickly and put that up, as well as about 100 other titles by the end of the summer of 2015 and from there, we dug in hard on that investment, and I’m happy to say that today, we have over 13,000 hours of audio described content on Netflix.
We’ve committed to creating audio description on all of our original content, and we also do our very best to make sure that we get the licensed content, we’ve actually added the need for audio description to our contractual obligations with our content partners.
It was so promising!
Even better, listening, reading and watching the feedback from the community through podcasts, blogs and YouTube was so gratifying.
You can’t under estimate the value AD brought to the community.
Fan 1:
AD changed my life.
Because I now had access, I find myself in conversations with more people because I can talk about the latest film or streaming series. I’m bonding with more people and I never realized this was possible.
Fan 2:
Watching AD, helped give me the courage to believe I could do more. I wanted to follow in their footsteps and today I’m working in a field that I never believed was possible.
Fan 3:
Today, I don’t watch anything unless it has AD.
With it’s new found popularity and all the opportunity that came from the CVAA, AD moved from one gig to the next.
Live television events like the Olympics and even award shows wanted AD.
More streaming services were coming online and sought out AD.
I was working constantly.
I was even invited on a cruise [laughs] only to find out it was a working trip.
In between projects, I had time to self reflect.
I listened to the community. That was eye opening.
No pun intended.
Ok, huh, AD has a sense of humor.
When I finally understood that my work wasn’t serving everyone equitably, I knew I needed to do a better job.
I realized, I had to better represent all cultures.
I even tapped into my ability to speak multiple languages and started bringing that to my work.
I was excited about not only my future, but what it could mean for my supporters.
Wow! At a time when I didn’t think the work could get better, it became even more of a passion!
— suspenseful music…
So where did things go wrong? What exactly happened to AD?
More after this message.
POD Access trailer
— a bouncy bass drum opens to bright synth trumpets announcing a fun, celebratory up-tempo Hip Hop beat
THOMAS:
If you’re currently thinking about launching a podcast, but have no idea how to start?
CHERYL: The types of podcasts, finances.
THOMAS: Recording, editing, marketing.
CHERYL: Social media, branding.
THOMAS: Theme music, a logo.
CHERYL: Maybe you’ve started one but feel as though you need a bit of coaching to really get it going? Where should we begin?
THOMAS: PodAccess.net.
CHERYL: In this series of 11 episodes, we’ll cover the basics of things to consider when starting a podcast.
THOMAS: Unlike most of the information you’ll find out there on the subject, we recognize the importance of accessibility.
CHERYL: Accessible content creation! The three ways you can potentially consume podcasts.
THOMAS: Listening, reading, watching.
CHERYL: We’ll share resources and practical tips from other Deaf and disabled podcasters in all stages of production. Tell a friend and an enemy to follow or subscribe to POD Access wherever you get your podcasts and on Instagram @PodAccess.
CHERYL: Connecting d/Deaf and disabled podcasters to audiences and each other!
— music fades out
Ten years after the CVAA, AD, was having even more of an impact on it’s supporters.
— Covid 19 Pandemic Announcement of Lock down.
In two thousand twenty the Covid 19 pandemic hit and forced corporations around the world to re-examine what work will look like in the near future.
In fact, for many it forced them to consider what workers look like.
Fortunately, some corporations had open minded people in charge and saw talent in the Blind community. Rather than just seeing fans of AD, or consumers,
— An uplifting synth opens to a pop beat.
they saw people who can help contribute and work with AD.
Wait, what?
You’re going to tell me I get to work directly with those who have been supporting me for years?
They say if you love your work, you’ll never work a day in your life.
— Music ends.
Well, I went on a permanent vacation when I got the chance to work with more of my supporters. It made me better!
— Dramatic bass heavy suspenseful drone.
And then things fell apart.
— Downton Abby Scene with TTS
Online Streaming companies turned to Artificial Intelligence (AI) and text To Speech (TTS).
For more on exactly why, we turn to a representative of Amazon
who further explained the company’s position during a panel discussion held in two thousand twenty two at the ACB convention.
Amazon guy:
We looked at the velocity of of how fast audio descriptions were being created. We looked at the very large back catalog of content. And we thought, you know, if we were going to really achieve the vision of what our audio description program is within Prime Video, that every single title is described, and, you know, audio descriptions are generally as available as closed captions are, then we would need to, you know, disrupt how audio descriptions are created. And so we’ve been working on text to speech audio descriptions for about a year and a half to date.
We have over 4000 text to speech tracks live on the service.
We have one person who kind of does everything the writing and the descriptions and the time coding, and then directs the text to speech engine. But we have every single track before it’s published on the service also goes through a human quality control review. So we are making sure that there is a second person involved in the process that reviews and detects those things that you know the writer may miss.
The more iterations you do on anything, the better quality product there will be.
Did anyone consider that yes, quality gets better over time, but in what other field do they release a sub par product to the consumer?
Who does that?
We spoke with AD supporters who shared their frustration.
Consumer 1:
I feel forced to make a decision every time I open a streaming service and launch a series or movie only to hear what I know is TTS.
I know they say they’re working to improve the quality of these voices, but it feels like they’re working harder to try and trick me.
Some of these companies using TTS narrators are crediting and naming the narrator.
Even if that TTS voice is based on a human, that doesn’t feel genuine to me. It feels like you’re just playing in my face.
Using AI in the production process goes beyond TTS.
The mixing is automated as well. And so it’s a separate process that is fully automated. You write the script and you get live feedback on the audio as you’re writing it, but then eventually, when it goes into mixing, that is automated as well.
We definitely realize there as well that is not as good as a human going through and mixing the track, but we think over time, we can build improvements into the track to get it as close or as to human quality mixing in the future.
The benefit of keeping everything automated, and that’s one of the reasons why we’re going fast, and we feel a little bit more comfortable about going quickly, is all of our tracks are never finished. I think that’s an important thing that we can do. We can very easily fix the script, and then, you know, automatically re voice and remix the track. If, if the improvements in text to speech come, we can, you know, use new text to speech engines to, you know, improve, improve the quality of the voicing. And as we’re, you know, making improvements in the automated mixing quality process as well. It’s very easy to for me to take all 4000 of those tracks and rerun them through the new, enhanced improvement in the mixing technology to improve the mix as well.
What?
So a consumer who watches a film that has poor quality description or doesn’t like the TTS voice, should just come back in a year or two to try again?
Who consumes content in that way? How could they say this is a working solution with a serious face?
As we hear from representatives from other streaming companies, it’s obvious, AI and TTS was always in play and part of the plan.
Peacock:
Certainly we are committed to innovation and would look to in the future, find ways of making audio description as efficient. Right and widely available as possible. And you know, text to speech could be one solution that that’s pursued.
I think you know there are ways that industry has to look to be able to deliver more content faster. And certainly, you know, technology solutions like text to speech may help.
Paramount:
Right now we don’t, and you know, we don’t have any immediate plans to do so, but I would say that, you know, my answer to that is, is similar to what Tom said, which is that
it is something that we’re looking at, and in part because, not as so much as a replacement for the live voice version of description, but as a way to possibly expand the amount of description that we have available on something like paramount.
The quality is improving there. It’s,
it’s certainly quicker than doing the traditional way. And it’s less expensive.
so I think it is something that we’re just looking at down the road as something that you know potentially could offer more content with description available.
Hulu:
In regard to Hulu originals, we currently only accept audio description that are read by human voice actors. So we don’t take auto generated or speech to text at this time, and we do try to make sure that we have consistency in our voice talent across episodes and seasons of a series
That was representatives from Peacock, Paramount and Hulu respectively.
Next, hear from HBO Max, Apple and Disney.
HBO Max:
We actually strive to really have our content voice acted. First and foremost, we’re a company. We’re an entertainment company. And you know, our value is really to create great entertainment. So it makes sense.
Apple:
We’re focusing on, you know, live human being, audio description. It’s kind of where we are at the moment. You know, I think, as some people have said, I think no one really knows where the future is going to go, but it’s not where we are currently putting our focus. We like the way things are at the moment.
Disney:
It is definitely something lucrative to look at. However. We have not ventured out that way just yet.

Consumer 2:
All of the effort these companies put into TTS, it’s like their trying to convince me that I want something I don’t want.
Today, it’s hard to find any platform that doesn’t employ TTS.
I think that’s a business decision that each studio needs to make.
We don’t want the trade off to be between text to speech and human that’s not the trade off we’re trying to address that. Really the trade off we’re trying to address is between no audio descriptions and text to speech.
— In order and panned from left to right:
A man confusedly says “Huh”, followed by a woman and dog with similar responses.
Consumer 3:
Why do, they, get to decide my options?
As a Blind consumer, I’m left with two choices, watch, or not.
Not much of a choice in my book!
The lack of choices can make some feel helpless. Not exactly what the CVAA was striving for.
Equal access, equal opportunity and equal respect for every American.
Equal access, equal opportunity, the freedom to make of our lives what we will; living up to these principles is an obligation we have as Americans.
What hurt me the most was, I expected someone to say something in my defense.
It was, after all, an organization of the Blind. I expected more pushback.
Not everyone was silent.
The community WAS speaking out against TTS and AI early on.
Online conversations like Blind Centered Chats, that often featured people from the community were vocal.
Clip from BCAD Chats 001
We didn’t talk about the fact that, well, how do we get the industry to really center blind people and blind and low vision people? Because right now I’m not sure if that is the case when it comes to audio description. I don’t always feel as though we are at the center of this. And there’s many reasons that I feel like that. Number one, I think this conversation about quantity and quality really does come down to who is being centered. Because when we talk about the quantity and really going for that, I think quantity, that whole, that kinda relates back to the whole compliance, let’s just get it done because the government is telling us we need to get it done. And that, to me-
NEFERTITI: Mmhmm. Checking that box.
THOMAS: Yeah. That, to me, brings about the Amazons, the AI, and all of that.
SCOTT N: [growls]
THOMAS: That’s what that’s about. And I think we all, the majority of us probably agree that that’s not really quality, you know, and we’re not centered in that conversation. That wasn’t about us. That wasn’t about bringing a good product to the people. That was more about, again, checking that box, like Nef said, and just making sure our numbers, and we do it efficiently, right? We do it on the cheap. That’s what that’s about. So, we’re not centered.
NEFERTITI: Do it on the cheap, do it at scale.
Other podcasts and content creators were vocal, perhaps not as loud as they could have been, but those in power either didn’t hear or care to hear the feedback.
This clip from an episode of Reid My Mind Radio’s Flipping the Script on Audio Description, reminds us that even audio description providers were sounding the alarm.
Clip from FTSAD: No One Will Save Us Part 2
When I invited Eric and Rhys to join me on the podcast,
I asked them each to bring three to five issues that most threaten the future of AD and some thoughts as to what we can do about them.
Eric:
I mean, there’s a few things that stand out.
Obviously, TTS synthetic speech, however you want to phrase it, I think that’s a problem.
It impacts our voiceover folks. But it also affects every area downstream for audio description.
A lot of companies that have no business being in this space are in this space. And they’re in it in major ways.
Rhys:
And it goes beyond just the sound of the voice, right?
Great narration track is often done by somebody who’s connecting to the material. Well, you know, who doesn’t connect to material? TTS doesn’t connect to the material, there’s no lived experience of being LGBTQ for a TTS voice. Whereas you get the human narrator, skilled, who’s doing content of that type, in the connection that they come through, it’s not performative, but it’s subtle, but it’s there, and it’s present. If you’re an adept listener of audio description, you can hear it, that person gives a crap about what they’re doing. And
we stand to lose all of that.
Does TTS serve a function in audio description? 100%. Like.
How to videos on YouTube, go for it.
That’s an entirely reasonable application for TTS. But if you’re taking premium content, what is it you’re trying to achieve by doing that?
The answer all your questions in life is money. That’s the old cliche.
It’s a pathetic thing that I have to say, but that’s really the answer. It’s always comes down to money.
— Sesame Street Word of the Day!
Can you say capitalism?
Ahh!The truth is, everywhere you look, there’s someone
trying to make a fast dollar by cutting costs and sacrificing quality.
Remember Economics class? Caveat Emptor or Let the buyer beware.
Consumers of all types should be educated enough to know the value of what they’re buying.
In this case, we’re talking about large companies, networks and streaming services that frankly have no experience with audio description.
So, how can they even begin to define quality AD?
Yet still they’re procuring millions of dollars in AD for old and new content.
They’re being misguided, misled because they’re talking to the clowns. They’re talking to people who are trying to sell them on something. They’re not necessarily talking to the audience.
I always encourage them, Are you sure, have you spoken to anybody about this? Has anybody told you by the way the audience doesn’t like this.
But the other part of this is an AD providers. If you’re confronted with that conversation, what do you do? Do you just go? Yes, sir, we can do it? Or do you go and take that opportunity to talk to them about what they’re asking for, to take that opportunity to go? Are you sure that the experience you want your audience to have of your content? That is your precious, precious item is a subpar experience? Because it doesn’t need to be.
Some how, some where, someone found a way to
either justify or convince streaming platforms and broadcasters that AD users want TTS.
TR in Conversation with Eric & Rhys:
Wow, so you telling me that companies come to you? And they say we want TTS?
Yep. They do.
AI and TTS definitely can help a single passion project based podcast to produce an idea in order to express himself and hopefully make a point. But even I agree, there are limits.

— Suspenseful cinematic music

AD is still working today, but they say their future just doesn’t shine as bright as it did in the early years following the CVAA.
I can only imagine the pain of living day after day, being turned down for gigs, while watching as AI and TTS slowly replaces you.
I’m sure that can really take a toll on someone.
it feels like the classic Hollywood story.
Rising up from the bottom after working so hard, and then reaching the top.
It’s one thing to lose after so much success, but I can’t even describe the pain of watching the community move on.
In fact, they even continue to celebrate as if their unaware of the losses.
Sources tell me, AD, is considering going underground.
The less value placed on them the less likely supporters will be inclined to pay for their services.
That’s not really an optimal solution for anyone, but I also understand it is a response.
If AD can only make their way into the hearts and minds of creators, chances are they can have even more of a future. Not only providing the equitable access but perhaps even more of a chance to flex their creative muscles.

AD:
For years, I based my value on my productivity and how I could impact the organization’s bottom line.
I no longer see it that way.
I’m aware of what I bring to the community.
Not only am I helping more people access content, but I’m impacting culture, education, creativity. I’m making relationships and connections that go beyond access.
Today, I’m working more with independent producers who know the value of human centered art.
We’re having conversations that go beyond sticking to the script. I’m being asked for my creative input. This gives me hope and is actually exciting!
A true crime story that ends with the presumed victim feeling optimistic?
This is far from the norm.
Strangely, in this story, the true victims, don’t seem to realize they’ve been lied to.
They don’t know they were deceived.
In fact, some are celebrating as if none of this ever happened.
If that isn’t a crime, tell me , what is it?

— Music continues and fades out.

TR:
Reid My Mind Radio Family!
A few episodes ago I told you all about Portraits and Portals.
A powerful art showcase uplifting disabled artists and disaster survivors of color .
The art takes on multiple formats including audio, poetry, song and visual art.
As this is Flipping the Script, I know we have listeners and readers who are quite skilled in writing description.
Well it just so happens that Portraits and Portals are asking the community to share some of their creative and poetic descriptions for the artwork posted on the site.
They’ll include it among the access resources available for visitors with vision loss.
Just go to PortraitsAndPortals.com and visit the artist pages.
You’ll be able to provide your description for visual art pieces.
Who better than those who rock with Flipping the Script.
I’ll link you there directly on this episodes blog post at ReidMyMind.com.

This concludes the Flipping the Script on Audio Description season.
I’ll be back and hope you stay tuned.

— Abbreviated “Your Official”

Official, uh uh official | Official, uh uh official !
Official, uh uh official | Official, uh uh official
You’re Official! | Reid My Mind Radio Family!

For all y’all thinking, hmmm, what can I do| I wrote a checklist so you can become official too
Sample: “One” Chuck D!
You just can’t be down if you don’t follow or subscribe | That’s the number 1 way to join this tribe!
Sample: “Two” Chuck D!
Shout it out! Tell all your friends and your foes | That’s the way we help each other and the podcast grows
Sample: “Three” Chuck D!
Buy merch, shirt or hoodie for yourself or your boo | Now look at that, you’re official too! Woo!

Official, uh uh official | Official, uh uh official !
Now whose official? | Official, uh uh official
You’re Official! Reid My Mind Radio Family!

Official, uh uh official | Official, uh uh official
I said who’s official? | Official, uh uh official
(Applause)
You’re Official! | Reid My Mind Radio Family!

— Music continues
Transcripts and more are available at ReidMyMind.com.
That’s my address on these internets.
And be careful, there’s a lot of artificial out here. It’s not all intelligent.
Sometimes it’s hard to know what’s real and what’s not.
There’s only one way to make sure you’re traveling safely out here, and you reach your intended destination.?
You gotta spell it right…
that’s R to the E I D!
“D” And that’s me in the place be!, Slick Rick

Like my last name.
— Reid My Mind Radio Outro
Peace

Hide the transcript

Blind Centered Audio Description Chat: – When AI Comes for the Blind

Wednesday, March 8th, 2023

While all the world is talking about AI, Artificial Intelligence, the Blind community has been dealing with what appears to be the eminent take over for the past few years. That’s the adoption of AI and Text to Speech in Audio Description.

In this last BCAD Chat of 2022 we wanted to discuss the pros and cons of AI and TTS voices narrating Audio Description.

Use this letter as a template to personalize and express your concerns about TTS in AD.

Shout out to Scott Blanks & Nefertiti Matos Olivares for the draft.

Join Us Live

The BCAD Live Chats can take place on a variety of platforms including Twitter and Linked In.

To stay up to date with the latest information and join us live follow:
* Nefertiti Matos Olivares
* [Cheryl Green]*(https://twitter.com/whoamitostopit)
* Thomas Reid](https://twitter.com/tsreid)

Listen

Transcript – Created By Cheryl Green

Show the transcript

Music begins
THOMAS: Welcome to the Blind-Centered Audio Description Chats. These are the edited recordings of the Blind-Centered Audio Description Live Chats!
CHERYL: The live is the most fun part! We get together, we start with a question, and then we invite up anybody from the audience who wants to come and chat with us, agree, disagree, shed light on something that we hadn’t thought about before, which is Nefertiti’s favorite. [electric whoosh]
NEFERTITI: I’m Nefertiti Matos Olivares, and I’m a bilingual professional voiceover artist who specializes in audio description narration! I’m also a fervent cultural access advocate and a community organizer.
CHERYL: I’m Cheryl Green, an access artist, audio describer and captioner.
THOMAS: And I’m Thomas Reid, host and producer Reid My Mind Radio, voice artist, audio description narrator, consultant, and advocate.
SCOTT B: Hi, I’m Scott Blanks. I’m a passionate advocate for the highest quality audio description in all of the arts. I’m the co-founder of the LinkedIn Audio Description Group and the Twitter AD community.
SCOTT N: Scott Nixon here. I’m an audio description consumer and advocate, hoping to be an audio description narrator very, very soon. [electronic whoosh]
THOMAS: Hey, Nef, why don’t you tell people how they could join the live recording?
NEFERTITI: That’s really simple. Just follow us on social media to keep up with important details, such as dates, times, and what platform will be using. On Twitter, I’m @NefMatOli. Cheryl?
CHERYL: I’m @WhoAmIToStopIt.
THOMAS: I’m @TSRied, you know, R to the E I D.
NEFERTITI: How about you, Scott?
SCOTT B: I’m @BlindConfucius. That’s Blind Confucius.
SCOTT N: And you can catch me on my social media, Twitter only. That’s @MisterBrokenEyes, Capital M r Capital Broken Capital E y e s.
[smartphone selection beeps]
CHERYL: Recording now!

NEFERTITI: Welcome, welcome. Welcome. And welcome, everyone! Tell a friend if you haven’t already. This is a conversation all about TTS, text to speech, and audio description. Place, the place for TTS in audio description. Is there one? What is it? Do you hate it? Do you love it? If you love it, we are particularly interested in hearing from you tonight. I’d love to hear from people who might change my mind or might make me think a little differently about this topic, because, frankly, I am strongly against TTS.
THOMAS: So, the conversation is all about AI and TTS, mainly TTS. But I figure we should have a little conversation about both, because they’re sort of used together. And I think there is a little bit of a difference when folks talk about AI, or artificial intelligence, in the audio description space and then TTS or text to speech. And so, a little bit of the difference: the AI, artificial intelligence, that usually refers to, so, that is computers sort of learning on their own and adjusting, making changes, and doing the things that humans would usually have to program. But the artificial intelligence and TTS sort of amounts to, if you, which I’m pretty sure everyone here has probably heard audio description when the speech comes in, and then the sound is ducked, well, that’s the artificial thing that’s happening right there. There’s sometimes when it’s done via AI. It’s not a human who’s actually sort of mixing the sound. The artificial intelligence is saying, “Okay, I’m gonna put it here. I’m gonna duck this down, go back up when the speech is finished. I’m gonna duck now. The speech is coming in, so I’m gonna duck the track.” And then you’re gonna hear mainly the speech. “And then when it’s finished saying the audio description, I’m gonna go back up with the track.” And so, the film itself will start playing louder. It’s not clean. It’s sort of a jumpy thing, kind of takes you out. And it’s kind of annoying. It’s kind of annoying. So, that’s part of the artificial intelligence.
There’s some AI that I think they’re also working on when it comes to, I don’t know if anyone’s doing that right now, but actually writing the audio description, as far as I heard. I think that might be in the works if it’s not actually out there. If someone knows if it’s been done, you tell me. But then the TTS part is what we all know as the text to speech or the synthesized speech. That’s the computer who is taking the job of the narrator. And I just mean that on that particular film. I’m not making a blanket statement about these things taking jobs from narrators, but, you know, in a way it’s happening. [laughs] So, yeah.
And so, the questions that we usually get into, the discussion usually is sort of like pro or con. Do you like it? Do you not? Are you okay with it? So, we can start off there. But that’s really, I don’t think that’s really where wanna to stay, because right now, whether we’re pro or con, I think we need to think about that the industry and those who are really offering this and pushing this well, they’re very pro. And they’re pro, we know, because not because of the artistic value of synthetic speech, but they’re pro because they wanna save some money, and as I like to say, so Jeff Bezos can go to space and whatever else he wants to do with all that money. That’s a whole nother conversation. I don’t know what you can do with all that money, but whoo. Anyway. But apparently what you cannot do is provide good audio description! [laughs] I said it!
I wanted to frame the conversation, but Neff and everyone else, Cheryl, Scotts, I’d say the two Scotts, if y’all wanna talk about pro/con, because I think the thing that would be interesting, maybe we can make the argument, maybe we can even invite some folks up to take a side of pro and con first and just to sort of get that to hear why people might actually be pro and hear what their arguments are, because it’s always good to hear from folks.
CHERYL: Cheryl here, and I will say that I’m con. I’m firmly on the con side that Thomas laid out. But I wanna be clear that that is not because I’m a professional audio describer, and I am sad that a computer is taking my job. And I could be, but I feel like, as in the sighted describer community, the narrator community, voice talent community, we need to be careful that our main argument against it isn’t, “I might lose my job.” It is awful to lose the job, but the point that is important to me is that my job is about creating audio description for the audience to have a wonderful, immersive experience. So, it’s the audio description and the user’s perspective, I think, that really should be paramount here when we discuss it. I think it’s fine to have a conversation about jobs, but that might be a different space because these conversations are blind-centered audio description conversations. So, I would ask that if there’s voice talent in here, that we keep it centered on what is the experience the audience is getting? And I just don’t feel like TTS offers, and especially the AI-written stuff, it doesn’t offer the nuance. I’ve seen things where the description is focused just on that moment between dialogue, but there was no opportunity to hear a description about anything that happened before the dialogue. There’s no context. These lines sort of float in space and don’t seem to connect and make a cohesive whole. So, I’ll stop there and hand it over to anybody else. Thanks.
NEFERTITI: How about you, Scott Blanks?
SCOTT B: I am one who is against TTS in the vast majority of the media that is currently being audio described. Let me elaborate. So, when I think about the arts or entertainment broadly writ, I’m thinking about film, TV, stage, other creative presentations, art, artistic exhibits, things like that. I feel in those spaces that as things currently stand that TTS, as has already been mentioned by a few people, is, it is not what is going to make the experience a quality one, and it doesn’t make it an accessible one. The point of audio description is accessibility, and audio description can also be considered an art form. But even if you just consider it on the accessibility side, if the accessibility tool is a synthetic voice that is mispronouncing words, that is, as has been mentioned, there’s an odd rhythm or arrhythmia to it that takes you out of the experience, then your experience is not only not as immersive, it’s not as accessible. And that’s the point of audio description in all of the contexts that we know it right now, and in a lot of the context where we don’t know it.
And I would say if I were, if I were to be pro audio description through TTS narration, it might be in some of those spaces where there is no option right now. If there was a way to access information that scrolled on a TV screen, real-time, newsworthy information, that might be something that I could see because having the quickest access possible to that information is really critical. And I don’t think it would be feasible to think that we could have a human standing by 24/7 on literally thousands of different networks, TV stations, feeds, whatever to provide that. But I think we have to kind of keep our focus. Most of the professionals here, the professionals on this panel that I’m fortunate to be alongside here are audio describing or writing for audio description or providing other contributions to the audio description field through arts and entertainment. And in that space, I don’t see that TTS has a place in the provision of audio description in 2022.
NEFERTITI: Beautifully said. I could not agree more. All right, Scott Nixon, let’s hear from you!
SCOTT N: All right. I would like to conduct a small thought exercise for the sighted people in the audience today. You’re at an art museum. You’re, well, you’re at the Louvre. Okay. You’re standing in front of the Mona Lisa itself, its glory, its majesty, its beauty. You’re drinking it in with your eyes. Imagine for a moment you couldn’t actually see the painting. You couldn’t experience it the way everybody else experiences it, so you have an audio description device plugged into your ear. Would you prefer a member of the artistic community talking passionately about the magnificent painting you’re seeing before you, [imitates stiff robotic voice] or would you like a robotic voice explaining to you what it looks like? [back to regular voice] That is what we’re talking about with TTS.
I myself am vehemently anti-TTS for audio description because it robs something you’re watching of its soul, okay? I have watched sitcoms and movies and various other forms of media with TTS audio description, as, you know, as a curiosity over the years, and it really does take something away from the experience. Why should we as a blind community have to have a lesser experience than everyone else just because a company wants to save a couple of thousand bucks? It is literally a matter of a couple of thousand bucks between bad TTS and even minimally good audio description. So, why not do it? The simple answer is they don’t think we matter enough.
So, at the end of the day, this is something I always say when I’m talking to people about accessibility, audio description, accessible websites, all that sort of stuff, “You are a company. You are ostensibly here to make money. If you make a quality product and an accessible product that vision-impaired and blind people will enjoy, we talk. We talk to each other. We tell people when something is good. If you build it, we will send you sack loads of money. So, why are you sitting on your butts doing something that you shouldn’t be doing?” And that’s me done for now.
THOMAS as Audio Editor:

THOMAS as Audio Editor:
Hey Y’all, I just need to interrupt for a moment.
During this live conversation, we had a challenge getting our technology to work. Well, that we is really me.
We wanted to play a clip in order to have a sample to discuss.

Mmy technology is working today so even though we didn’t have the chance to discuss it, you can have a chance to hear the sample.

Check this out!

Downton Abbey clip:

Test to Speech Audio Description Narrator:
In the English countryside, a turn of the century train barrels past the lake

it rumbles by dead leaves and bare branch trees. puffs of white steam ripple out from its engine and below on to the rolling green hills.

On board a large black haired man in his late 40s peers out his window. Steam envelops the wires of utility poles.

In a village, a wire travels between quaint stone houses to a telegraph office.

NEFERTITI: Amazon Prime is where you can find this example. It’s a show that was hugely popular called Downton Abbey from our British neighbors over there across the pond. What isn’t beautiful about this experience is that as majestic as the show is, it has TTS. And the TTS says a lot of the things or is guilty of a lot of the things that Scott Blanks mentioned: mispronouncing names, misnaming names. So, in addition to the audio description script being kind of crappy, then on top of that, you have this robotic voice who, for those of you who are blind and in the audience who use a screen reader, it’s worse than like the Eloquence screen reader. Eloquence, for those who are not aware, is the most popular, widely used screen reader that blind people use to get around on the Internet, on PCs, on Windows machines. So, I mean, it’s just super distracting, really kind of offensive, and just not at all in keeping with the content, which is very dramatic and passionate! But then you have this [imitates robotic voice] TTS voice: The train rides down the rails. You know, it’s just, it’s, and not for nothing, but I sound great compared to the TTS just now. So, it’s just, it’s just so inappropriate.
SCOTT B: It’s Scott Blanks just to jump in. And if you’ve not ever enabled audio description on Prime video, you can do that once you start playback of an item. There should be an audio and subtitles option on your playback screen that you can access, and in there, you would wanna choose, in the case of Downton Abbey, English audio description.
SCOTT N: Just as an example of bad versus good, the American sitcom, The Big Bang Theory, huge hit in its day. Audio description turned up on Amazon Prime here in Australia about a year ago, and I was all gung-ho and ready to listen to it. I put on the first episode, bang. TTS. Completely robbed the show of its humor and its charm. I gave up after two episodes. Now, this year, HBO Max in America have apparently provided a human-narrated audio description track, and I was played a brief sample of it: 12,000% improvement. It gave the humor of the show. The audio description narrator was playing along with the jokes, smiling at the right times, frowning at the right times with his voice and all that sort of stuff. And it really did enhance the experience. So, the difference between TTS and human AD is like night and day. It’s just really a really important thing. And like I said, it helped to bring the soul of the show alive to the people who can’t see the soul that they put up on the screen. And that’s me.
THOMAS: In this conversation, TTS is sort of the demon. It’s the bad guy. But, you know, the technology’s not the bad guy. Like, we use TTS as blind people, as people with disabilities in general. We use TTS. TTS can, I love my screen reader. It gives me access. The screen reader is my input. It’s the way that I take in information. That is, that’s my access. The screen reader, that’s my guy! [laughs] Like, you know what I mean? Because he’s helping me out all the time. And then in order for me to have digital output, screen reader’s my guy. Like, I need him or her or them, right? And so, it shouldn’t necessarily be demonized. And I think that sometimes there are other people with other disabilities that make use of access technology, of TTS as well. So, Cheryl, you wanna talk about that?
CHERYL: Thanks, Thomas. I feel like you’ve framed it up so beautifully. The point that I wanted to make is that I have listened to different panels and read things and heard people arguing against TTS, which again, to reiterate, [chortles] I am not for TTS, especially as the Scotts pointed out, in a museum or a work of, a film, an art piece, an art film. But what troubles me is sometimes the reasons given end up incorporating a lot of ableist slurs and a lot of really harsh language, which I’ve heard none of tonight. But what I want us to be careful is, like Thomas said, to not demonize the technology. And for those folks who have a lot of communication through one of these systems where they’re typing or selecting images and some kind, a synthesized voice comes out, that’s communication. And so, it’s not the voice that’s “awful and soulless and inhuman.” I just want us to be careful. And when you leave this session and you go out and you promote or you speak about the harms and the problems with TTS, that you be careful to not be too ableist and throw augmentative and alternative communication users under the bus while insulting the sounds of these voices. It’s not the sound of the voice, it’s the application. And like Thomas was talking about, when the AI adjusts the volume of the soundtrack for this TTS to come in, it is like my head starts spinning. It’s just so jumpy. It’s so, it’s not artistic, and it doesn’t fit the vision of the film or the show. So, I’ll pause there.
THOMAS: And I also wanted to jump in with two podcasts ‘cause I think Cheryl, you had a podcast with a AAC user, and I think, so, if folks wanna kind of get to see how people use these devices and how it’s so intertwined with their life, that’s one. So, what’s the name of that podcast, Cheryl? I think it’s called Pigeonhole.
CHERYL: Oh, my! No, no. People should go to endever’s podcast. AAC Town is the podcast that endever* and their comrade, Sam, run. They’re both AAC users full-time or nearly full-time, and they have a podcast. It’s all transcribed. But yes, I did have endever* on my show, Pigeonhole, one time.
THOMAS: That’s what I was talking about.
CHERYL: Yeah.
THOMAS: You had them on your show.
CHERYL: But it was to talk about—
THOMAS: Let me do what I gotta do, Cheryl! [laughs]
CHERYL: OK! [laughs]
THOMAS: You always shout out my podcast, and so I wanna shout out yours. But mainly because it applies, right? I don’t wanna just shout, you know, I’m not just randomly shouting out podcasts. Although I do that around here. Every two hours I open up my window and I shout it out, “Pigeonhole!”
BOTH: [guffaw]
The other one is I was gonna say, now I am gonna do a promo of mine, is because I had a conversation with Lateef McLeod. And Lateef McLeod is a AAC user. And in that episode, we really go into some of the other issues around TTS that never necessarily get talked about. Lateef McLeod is an African American, and the voices that he had all his lives don’t really represent him until he got a voice that was a synthesized voice of a Black man. And so, you know, these issues are big, right? We always talk about it, like, these issues are really big. And so, and in that, Nefertiti was actually in that episode, too, where we did a little bit of a little skit about TTS that touches on a bunch of these things. So, anyway, it was a cool episode, I think. And so, both of those, check it out, and we get into these conversations as well. That’s it. I’m Thomas. I’m done.
NEFERTITI: Yeah. So, I think the summary here is let’s express ourselves, but be mindful to not sort of turn around while advocating for one accessibility, mm, putting down another, you know, or minimizing, punching down on another. So, I think that’s a great point. And we had a great clip to show you related to that, where human narration meets TTS and how it was used judiciously, minimally, but in a way that really drove home the point of where maybe it’s appropriate.
Scream Trailer from Social Audio Description Collective

AD Narrator – Nefertiti Matos Olivares:
The lights are on In a white suburban house at night. A silver cordless landline rings with the ID, Unknown Name.

In the kitchen, Tara pushes reject on the cordless while holding her smartphone. She is a thin light skinned Latina teen with long wavy dark hair pulled back in a ponytail.

She’s just texted

TTS Receiving Text Message:
Mom’s out of town again you should come over here. Free dinner, Many binge watch options.
AD Narrator – Nefertiti Matos Olivares:
Amber responds.
TTS Sending Text Message:
Have to do better.

TTS Receiving Text Message:
Unlock liquor cabinet.

— Landline phone rings

TTS Receiving Text Message:
You should answer it

TTS Sending Text Message:
How did you know my landline was ringing?
Amber?

TTS Receiving Text Message:
This isn’t Amber.

Tara speaking on landline:
This isn’t funny amber

Deep Menacing Voice over landline:
Would you like to play a game? Tara?

Suspenseful Crescendo closes the scene.

THOMAS: That had some really different reactions that I wonder where people stand with.
NEFERTITI: First, I wanna say that this is for a Scream trailer that the Social Audio Description Collective described. It’s a bit of a hacker horror type film. And we had a human narrating the audio description, but there was a scene between somebody who was on camera and somebody who was off camera, and they were texting one another. And so, we decided, full disclosure, I’m part of the Social Audio Description Collective, we decided that why not use a synth to say what those lines of texts were rather than having the human describer say them? Just like we blind people experience TTS all the time with our screen readers, etc., why not just put one of those voices to those text messages? And it was a very brief exchange but still sort of drove home the point.
CHERYL: It worked so beautifully because I’m watching the screen, and I’m seeing basically a computer screen, words pop up on a computer screen. So, hearing that screen reader voice read it was really cool. And it really uplifts, in my mind, the ingenuity and creativity of disability community. Like, who would’ve thought to do that besides people who interact with these voices all the time? I thought it was such an add to the, it elevated the art, I thought.
THOMAS: Yeah, and I think I remember that there were some comments from folks who I don’t think were blind who were very negative toward that text to speech being included in there. And it was like, wow, like this is totally my experience. This is a text message. That was what a text message sounds like to us.
NEFERTITI: Mmhmm.
THOMAS: You know? And so, again, to me, highlighting that no, audio description should always center blind people. And so, blind people need to be a part of this, and blind people need to be a part of that conversation, which is part of my issue personally with, and so, advancing this a little bit, is the framing of this conversation of audio description and the way it’s been framed within the community by those outside of the community, those creating it, the corporations, right? Is that hey, TTS is good because you will get more. So, it’s either, if you want more audio description, then you take TTS. And that, framing TTS that way, is the biggest problem that I have with this entire subject is because we are being told we are being given options, and it’s two options, and we have never been consulted. And if they tell me that, “Oh, no, you were consulted as a community because we issued a survey that some folks got to fill out,” I don’t care about that. Because the thing is, is that it’s still based on that option that you give me. So, a lot of people would say, “Well, if these are my only two options, TTS or no audio description,” I can see why a lot of people would go that route.
NEFERTITI: Mmhmm.
THOMAS: But that’s not the route, that’s not the choice that we should be given. Why are you giving us those two choices? Those aren’t really even choices. And so, that’s my, really, my biggest problem with this whole conversation. I think history sort of says that when large corporations get their hands on something and have it in their cold hands and their cold hearts [chuckles] to do something and to get it done and to save a penny or two, they’re gonna do it. It’s gonna happen. And so, right now, my concern is that is this conversation about pro or con, does it even matter at this point? Is this inevitable?
And so, should the conversation actually move into something else like, “Hey, Amazon, hey, you corporations, why don’t you, you should, you need to be including us in this conversation”? Because like Scott, I think Scott B., you mentioned some other opportunities where, you know, okay, wait. Text to speech, I’ll take it here. This would work. This would help my life here in this particular case. And I’m wondering if there are other examples of that, that apply to film. Can we talk about either this framing of no, of more AD with text to speech or not, but also, is this inevitable? Do y’all think it’s inevitable? Do we still have a chance to say no? Or should we be talking about, hey, let’s come to the “negotiation table” and have these conversations and find out where the blind community says, “Okay, this would work for us”? I wanna throw that out there.
NEFERTITI: Mmhmm. And the blind community and our allies. Let’s never underestimate the power of allyship and togetherness. You know, this is accessibility.
THOMAS: Absolutely.
NEFERTITI: But it’s not to exclude our sighted allies. We center blind people here, but we are here, and we want to be part of the conversation just as much as our sighted allies have been already.
But I would like to hear a little bit from Scott Blanks about this idea that I’m sure is not exclusively something that, you know, just sort of a light bulb went off in his head, but something that he has taken and done something about. And it’s all about advocacy and a campaign of sorts. Because, Thomas, what you were saying about so, what are our choices? No description or TTS? And is this sort of the end of the road? Do we just let them do them being the cold-hearted, cold-handed, as you put it, companies out there to save a penny go the way of TTS, or do we do something about it? Can we do something about it? And I think that Scott has come up with a way that we can, if people get behind it. Scott Blanks, do you wanna talk a little bit about that?
SCOTT B: I have found that there are a lot of different ways to engage with companies, not just talking about audio description, but in so many different things. And particularly if you’re a person with a disability and unfortunately, you have to fight for a lot of stuff, small things, large things. Sometimes small things feel big. And there’s more of that than there should be. That’s a different topic, though. But I find that engaging with companies, it’s very easy to do that in places like social media and in sort of those public spaces. But what tends to be a little bit more of a lift for us, but also, I think has more impact on these companies, is when you start writing to them directly and when they start hearing from people in numbers.
So, one of the things that I did a few months ago was I took a run at a very basic, it’s sort of a template of sorts, a very, very rudimentary template that someone can take and use however they would like to reach out to, if they know of an entity who is providing TTS audio description, and they would like to talk about why they feel like that company should look at doing it a different way. This is a, it’s in a Google document that anyone can access. I would say the best thing you can do is connect to the LinkedIn audio description group, the Twitter community, or you can come find me on LinkedIn or anyone here really would probably be able to get you access to that link. It’s a public link, and it is available for anyone to view, copy, and do with as you see fit. But I believe it’s important. If companies don’t hear from us, and they’re doing a thing, then they think they’re doing that thing correctly. They believe that that’s how it should be done unless they start hearing from people.
And listen, I’m not under any sort of illusion that writing a bunch of letters is guaranteed to make a change. But I don’t like the idea of something becoming so rooted in, and the expectation is that this will be the way things are for now and evermore and thinking that we didn’t try hard enough. And I believe that part of advocacy is it’s not as flashy, but it’s getting those letters written. It’s getting that contact to these companies. And all of these c
THOMAS: Cool. Well, that concludes this week’s conversation. Why don’t y’all keep the conversation going on social media.
CHERYL: Use #ADFUBU, for us by us, #DescribeEverything, and #AudioDescription.
NEFERTITI: And hey, you know we’re out here, right? Mmhmm! Gathered and galvanized y’all. If you haven’t joined us yet, what are you waiting for?! You can find us in the LinkedIn Audio Description group and the AD Twitter community. We know that your participation will only make these spaces better.
Music fades out!

Hide the transcript

Microsoft Seeing AI – Real & Funky

Wednesday, August 2nd, 2017

!T.Reid wearing a hat with a "T" while the Seeing AI logo is imposed on his shades!
Okay, I don’t usually do reviews, but why not go for it! All I can tell you is I did it my way; that’s all I can do!
It took a toll on me… entering my dreams…
I’m going to go out on a limb and say I have the first podcast to include an Audio Described dream! So let’s get it… hit play and don’t forget to subscribe and tell a friend to do the same.

Listen

Resources:

Transcript

Show the transcript

TR:

Wasup good people!
Today I am bringing you a first of sorts, a review of an app…

I was asked to do a piece on Microsoft’s new app called Seeing AI.for Gatewave Radio.

The interesting thing about producing a tech related review for Gatewave is that the Gatewave audience most likely doesn’t use smart phones and maybe even the internet. However, they should have a chance to learn about how this technology is impacting the lives of people with vision loss. Chances are they won’t learn about these things through any mainstream media so… I took a shot… And if there’s anything I am trying to get across with the stories and people I profile
it’s we’re all better off when we take a shot and not just accept the status quo

[Audio from Star Trek’s Next Generation… Captain La Forge fire’s at a chasing craft. Ends with crew mate exclaiming… Got em!]
[Audio: Reid My Mind Radio theme Music]

[Audio: Geordi La Forge from Star Trek talk to crew from enemy craft…]
TR:
Geordi La Forge from Star Trek’s Next Generation , played by LeVar Burton, was blind. However, through the use of a visor he was able to see far more than the average person.

While this made for a great story line, it also permanently sealed LeVar Burton and his Star Trek character as the default reference for any new technology that proposes to give “sight” to the blind.

[Audio: from intro above ending with Geordi saying…
“If you succeed, countless lives will be affected”
TR:
What exactly though, is sight?

We know that light is passed through the eye and that information is sent to the brain where it is interpreted and
quickly established to represent shapes, colors, objects and people.

A working set of eyes, optic nerves and brain are a formidable technological team.
They get the job done with maximum efficiency

Today, , with computer processing power growing exponentially and devices getting smaller the idea that devices like smart phones could serve as an alternative input for eyes is less science fiction and well, easier to see.

There are several applications available that bring useful functionality to the smart phone ;
* OCR or optical character recognition which allows a person to take a picture of text and have it read back using text to speech
* Product scanning – makes use of the camera and bar codes which are read and the information is spoken aloud again, using text to speech
* Adding artificial intelligence to the mix we’re seeing facial and object recognition being introduced.

Microsoft has recently jumped into the seeing business, with their new iOS app called Seeing AI… as in Artificial Intelligence!
There’s no magic or anything artificial about these results, they’re real!

In this application, the functionality like reading a document or recognizing a products bar code are split into channels. The inclusion of multiple channels in one application is already a plus for the user. Eliminating the need to open multiple apps.

Let’s start with reading documents.

For those who may have once had access to that super-fast computer interface called eyes , you’re probably familiar with the frustration of the lost ability to quickly scan a document with a glance and make a quick decision.

Maybe;
* You’re looking for a specific envelope or folder.
* you want to quickly grab that canned good or seasoning from the cabinet.

With other reading applications you have to go through the process of taking a picture and hoping you’re on the print side of the envelope or can. After you line it up and take the picture you find out the lighting wasn’t right so you have to do it again.

Using Microsoft’s Seeing AI you simply point the phones camera in the direction of the text

[Audio App in process]

Once it sees text, it starts reading it back! The quick information can be just enough for you to determine what you’re looking for. In fact, during the production of this review, I had a real life use case for the app.

My wife reminded me that I was contacted for Jury duty and I needed to follow up as indicated in the letter. The letter stated I would need to visit a specific website to complete the process. I forgot to put the letter in a separate area in order to scan it later and read the rest of the details. So rather than asking someone to help me find the letter, I grabbed the pile of mail from the table and took out my iPhone.

I passed some of my other blindness apps and launched Microsoft Seeing AI. I simply pointed the camera at each individual piece of paper until finding the specific sheet I was seeking. The process was a breeze. In fact, it was easier than asking someone to help me find the form. Ladies and gentlemen, that’s glancing!

Now that I found the right letter, I could easily get additional information from the sheet by scanning the entire document. I don’t need to open a separate app, I can simply switch to a different channel, by performing the flick up gesture.

Similar to a sighted person navigating the iPhone’s touch screen interface , anyone can non visually accomplish the same tasks using a set of different gestures designed to work with Voice Over, the built in screen reader that reads aloud information presented on the screen.

Using the document channel I can now take a picture of the letter and have it read back.

One of the best ways to do this is to place the camera directly on the sheet in the middle and slowly pull up as the edges come into view. I like to pull my elbows toward the left and right edges to orient myself to the page. Forming a triangle with my phone at the top center. The app informs you if the edges are in view or not.
Once it likes the positioning of the camera and the document is in view, it lets you know it’s processing.

[Audio: Melodic sound of Seeing AI’s processing jingle]

You don’t even have to hit the take picture button. However, if you are struggling to get the full document into view ,
you could take the picture and let it process. It may be good enough for giving you the information you’re seeking.

If you have multiple sheets to read, simply repeat.

Another cool feature here is the ability to share the scanned text with other applications. That jury duty letter, I saved it to a new file on my Drop Box enabling me to access it again from anywhere without having to scan the original letter

Let’s try using the app to identify some random items from my own pantry.

To do this, I switch the channel to products.

[Audio: Seeing App processing an item from my pantry…]

What you hear, is the actual time it took to “see” the product. All I’m doing is moving the item in order to locate the bar code.
As the beeps get faster I know I am getting closer. When the full bar code is in range, the app automatically takes the picture and begins processing.

[Audio: Seeing AI announces the result of the bar code scan… “Goya Salad Olives”

It’s pretty clear to see how this would be used at home, in the work environment and more.

Now let’s check out the A I or artificial intelligence in this application.

By artificial intelligence, the machine is going to use its ability to compute and validate certain factors in order to provide the user with information.

First, I’ll skip to the channel labeled Scene Beta…
Beta is another term for almost ready for prime time. So, if it doesn’t work, hey,, it’s beta!

Take a picture of a scene and the built in artificial intelligence will do its best to provide you with the information enabling you to understand something about that scene.

[Seeing AI reports a living room with a fireplace.]

This could be helpful in cases like
If a child or someone is asleep on the couch.

[Audio: Action Movie sound design]

I can even picture a movie starring me of course, where I play a radio producer who is being sought by the mob. The final scene I use my handy app to see the hitman approaching me. I do a round house kick…
ok, sorry I get a little carried away at the possibilities.

While no technology can replace good mobility travel skills I can imagine a day where the scene identification function will provide additional information about one’s surroundings.
Making it another mobility tool for people who are blind or visually impaired.

Now for my final act… oh wait it’s not magic remember!

Microsoft Seeing AI Offers facial recognition.
That’s right, point your camera at someone and it should tell you who that person is… Well, of course you have to first train the app.

To do this we have to first go into the menu and choose facial recognition.
To add a new person we choose the Add button.
In order to train Seeing AI you have to take three pictures of the person.
We elected to do different facial expressions like a smile, sad and no expression.
Microsoft recommends you let sighted family and friends take their own picture to get a good quality pic.

The setup requirement, while understandable at this point sort of reduces that sci fi feel.

After Seeing AI is trained, once you are in the people channel
when pointing your camera in the direction of the persons face, it can recognize and tell you the person is in the room.

[Audio: Seeing AI announces Raven about 5 feet in front.]

Seeing AI does a better job recognizing my daughter Raven when she smiles. That too me is not artificial intelligence because we all love her smile!

The application isn’t perfect. it struggled a bit with creased labels, making it difficult to read the bar code.

Not all bar codes are in the database. It would be great if users could submit new products for future use.

As a first version launch with the quick processing, Seeing AI really gives me something to keep an eye on. Or maybe I should say AI on!

Peering into the future I can see;

* Faster processing power that makes recognition super quick,
* Interfacing with social media profiles to automatically recognize faces and access information from people in your network
* lenses that can go into any set of glasses sending the information directly to the application not requiring the user to point their phone
at an item or person and privately receiving the information via wireless headset.
That could greatly open up the use cases.

In fact, interfacing with glasses is apparently already in development and
the team includes a lead programmer who is blind.

Microsoft says a Currency identification channel is coming in the future;
making Seeing AI a go to app for almost anything we need to see!

The Microsoft Seeing AI app is available from the Apple App store for Free 99. Yes, it’s free!

I’m Thomas Reid
[Audio: As in artificial intelligence!]
For Gatewave Radio, audio for independent living!

[Audio: Voice of Siri in Voice Over mode announcing “More”]

I don’t know if that’s considered a review in the traditional sense, but honestly I am not trying to be traditional.

The thing is, thinking about the application started to extend past the time when I was working on the piece…

That little jingle sound the app makes when it’s processing… it started to seep into my dreams…
[Audio: Dream Harp]

[Audio: “Funky Microsoft Seeing AI” An original T.Reid Production]

The song is based around the processing tone used in the app with the below lyrics.

(Audio description included in parens)

(Scene opens with Thomas asleep in bed with a dream cloud above his head)

The processing sound becomes a sound with Claps…

(We see a darkened stage)

(As the chorus is about to begin spotlight shines on Thomas & the band)

Chorus:
Microsoft Seeing AI
Helping people see without their eyes

Microsoft Seeing AI
Helping people see without their eyes

(Thomas rips off his shirt!)

Verse:
Download the app on my iPhone

{Background sings… “Download it, Download it!}

Checking out things all around my home

(Thomas dances on stage)

Point the camera from the front
Huh!
Point the camera from the back!

I’m like;
what’s that , what’s this
Jump back give my phone a kiss!
Hey! (James Brown style yell!)

(Thomas spins and drops into a split)

Chorus:
Microsoft Seeing AI
Helping people see without their eyes

Microsoft Seeing AI
Helping people see without their eyes

(Back in the bed we see Thomas with a fading dream cloud above his head)

Ends with the app’s processing sound.

TR:
Wow, definitely time to move on to the next episode…

With that said, make sure you Subscribe wherever you get your podcasts. Tell a friend to do the same – I have some interesting things coming up I think you’re going to like.
And something you may have not expected!

[Audio: RMMRadio Outro]
TR:
Peace!

Hide the transcript