Cutting Through the BS of Corporate Video: Episode 3 – How Does the Technology Help?

Here we are again with Episode 3 in our series on “Cutting Through the BS of Corporate Video“!

In this episode, together with Jeff Sengpiehl, The Post Doctor, we examine the role of technology/functions in two specific areas within the Content Factory – AI and “Findability”. The primer can be a tool for good or bad, but there are limits and prerequisites one must be aware of. The latter is the fundamental business need that AI and the rest of the tech stack must serve.

Watch the episode here (don’t forget to share, like, subscribe, etc.!):

Join the discussion on Varde’s LinkedIn page and stay tuned for the upcoming episodes where we also bring in manufacturers and end-customers!

Takeaways

  • AI in the media supply chain is not new. Facial recognition, speech to text, and object recognition have been in production workflows for over a decade. What is new is the scale and the expectation.
  • Hallucination risk is highest exactly where the output looks most confident. In video that shows up as wrong names on faces, invented attributions, or synthetic elements nobody approved.
  • The line between acceptable AI use and unacceptable AI use is supervision, not the technology itself. Generative fill on a background is a different ask than an AI-generated CEO delivering a message.
  • Storage is not the deliverable. Retrieval is. An organization can build a content factory and still fail if nobody can find what is in it.
  • Metadata has three layers with three different lifespans: production, operational, and preservation. Most systems collapse them into one and lose two.
  • The real test of a system is whether someone with no training in it can find a specific piece of content in five minutes without calling anyone.

Full transcript

Lightly edited for readability. Filler words and false starts removed.

Introductions (0:07)

Jeff Sengpiehl: Jeff Sengpiehl. I’ve got decades in media technology, starting back at ABC and going all the way into AI. Along the way, broadcast engineering, post-production, systems integration, facilities build-out. I’ve had CTO and high-level engineering roles across post and storage: Light Iron, Chainsaw, KeyCode, Qualstar. Today I’m working as a consultant and fractional CTO as The Post Doctor. It’s vendor neutral, focused on production, post, infrastructure, live storage, asset management, expanding markets, and workflow integration. I’ve been involved with SMPTE, HPA, and SBE. I also publish and podcast under The Post Doctor.

Eivind Sandstrand: And I’m the founder and principal at Varde Media Solutions. Varde is a new, small, but growing little boutique consulting outfit. We’re also a solution reseller and a solution builder, as opposed to a traditional systems integrator. What’s unique about us is that we focus exclusively on non-broadcast clients. What we do is help our customers build the kind of media operation they truly need in order to meet the actual business objectives of creating, managing, and publishing the content, without dragging them down that rabbit hole of traditional broadcast engineering. And those of you who know what I mean by that, you know what I mean by that.

AI in the media supply chain (1:29)

Jeff: Shift gears back to AI. We’ve seen tons of presentations done in PowerPoint about what AI is and what that means in the marketing deck. So what do you feel it actually means in the media supply chain?

Eivind: It can do a lot of good. It can also do a lot of not so good things. I was having a very interesting conversation with some younger folks the other day, this generation around 20 to 25. They’re not exactly in favor of AI, and these happen to be two creative types. But we kind of landed on this: there’s AI for good, there’s AI for evil, there’s AI for stupid, and there’s AI for profit.

AI in the media supply chain has really existed for many years already. It’s got to be close to 15 years since I first saw facial recognition inside a MAM system. It was very rudimentary, but it worked. We’ve been doing speech to text for a long time for automatic subtitling and captioning. We’ve been doing object recognition for quite some time too. All of this is about generating metadata, about understanding the content we actually have. A friend of mine calls it content intelligence, and I think that’s a good term for it.

But there are still legitimate concerns about AI in content and in communication. For AI to be optimally useful: first, put the human in the loop. Second, your content has to be online — you can’t lock it away in old systems, you have to make it accessible to these AI engines, whether on premises, private cloud, or public cloud. You need access to proxy files, you don’t need to read the big ones. You need metadata standards, normalized schemas. You need expectations around what error rate is acceptable in subtitling and captioning — the ADA actually sets requirements around this. And you need governance decisions on what services to use, plus the human in the loop, plus workflows for quality assurance.

Jeff: The interesting thing is every one of those blockers you mentioned is a project to be undertaken, and none of them are projects that can be done by AI. That’s the part people don’t think about. If you’re going to get value from AI in video, you’ve got to do all this boring prep work first. Once you’ve got the domain and expertise set, then AI can come in and do the job you want it to do and give you close to the results you want. AI’s never going to be perfect, I don’t believe.

Eivind: Yeah. From there, once you set this up the right way, it becomes a different level of decision. How far do we want to take this? The AI field is evolving so rapidly that who knows what we’re going to be able to do tomorrow if we just enable it today.

Should anyone use AI to create professional content? (5:14)

Eivind: So do you think anyone should use AI to create professional content?

Jeff: This is the third rail of creative today, this discussion. I think the answer really hinges on what “create” means. AI is really good at time-consuming and tedious tasks — that works great inside the content factory. Transcription, object identification, facial identification, logging, format prep, even getting into first pass.

Then there’s generative AI, and that really depends on your goal. If your goal is to create an artificial human and have them deliver the message in place of the CEO, that’s a very sticky wicket. If you shot the CEO and there’s a blank space next to them and you use generative fill in Adobe Premiere to fill in the background, that’s a different ask. Getting rid of time-consuming and tedious tasks is a great use for AI. But I’m not at the point where I’d say it can be fully trusted. It’s going to need human supervision, even as it gets better.

One thing I saw this year was an animated feature done with AI where each character had their own person iterating on it, as much as a standard animator would. So I don’t think we’re quite there yet.

That also raises the hallucination question. I think the enterprise folks treat that as a rounding error. Where do you feel hallucinations drop in?

Eivind: It’s a very real thing. We see it every day in our use of different AI tools. I had to instruct Claude to stop imagining things, and even then I still can’t fully trust it. Hallucination is a real risk, not a theoretical one, and that risk is highest exactly where the output looks most confident. The more that’s at stake, the higher the risk. In the context of video, hallucinations can show up in generated descriptions, incorrect transcripts, wrong attributions, wrong names on faces, synthetic elements you never approved. Someone’s voice you don’t own could end up in a video you publish. Suddenly Tom Cruise is there and he’s got enough money to sue you. A competitor’s logo could get inserted, or the wrong product packaging. Those are bad outcomes for any organization, educational or financial services.

Would you let a person with six fingers, or a person who occasionally uses foul language, be the representative of your brand? No, you wouldn’t.

Eivind Sandstrand

The illustrated question is: would you let a person with six fingers, or a person who occasionally uses foul language, be the representative of your brand? No, you wouldn’t. That’s what unsupervised generative AI can do for you.

Jeff: Yep. The key thing is the right AI, used for the right task, inside the right process, is still going to deliver a real benefit. But it’s not just inspecting at the end — it’s inspecting every part of the process to make sure it’s on track and doing what you need it to do.

Eivind: So the question then, just like with transcription, we’ve raised the correctness level. Ten years ago it was around 89%, now we’re at 95, 96, 97. Are we ever going to hit 100%? And second, should we even then believe it?

Jeff: I don’t think we’re ever going to get to 100%, because we’ll always have people who talk in ways that don’t parse cleanly. But I think other vectors will come into play. If I’ve got an accent, or a lisp, or a raspy voice, the AI audio transcription may not be perfect, but you could pair it with AI lip-reading transcription that compares the two and says, no, that’s not the right word, this is what’s intended. As it gets used to different voices and builds a content library of what you do, it’ll get closer and closer. Even humans aren’t 100% on that.

Where’s the line: dubbing, lip generation, and audience pushback (10:47)

Eivind: A couple of years ago somebody was demonstrating, I think at NAB, how they’d taken a CEO-type character, done the speech-to-text thing, translated it, and then done not lip reading but lip generation. Hugely beneficial for a large international company. But where do we draw the line? Like you said, filling in a keyframe next to someone is one thing. How far across the line is something like that?

Jeff: That’s an interesting question. A year or two ago there was a production brought to the screen by AMC Theaters, produced in Sweden. The actors did an ADR read — automatic dialogue replacement, for those who don’t know — in English, and then they used AI to change their own mouths so it looked like they were speaking English. The voice was actually them; they were just changing their lips. I think we’ll slowly begin to spread that out.

Both of these use cases are highly supervised. We don’t have people saying new and different things that aren’t part of the script or that go against the company’s or filmmaker’s vision. So I think we’ll get there slowly. But at the same time there’s pushback from younger viewers who catch anything they think is AI and check out. So the audience is going to demand a certain level of production from enterprise video, and from video in general, and that will determine how it moves. It may not move the way we think based on the technology, because we’re technologists, we build things, and then people come along and say, yeah, but I want to use it like this.

Findability: storage is not the deliverable (13:21)

Jeff: So let’s say we’ve done it. We’ve built the content factory, we’ve put all our video in it. What happens if nobody can find anything?

Eivind: Something has gone very wrong in the design and construction of that content factory. Storage is not the deliverable here. Retrieval is. Without retrieval, without finding stuff, you can’t use it, you can’t do anything, you just keep making more.

Storage is not the deliverable here. Retrieval is.

Eivind Sandstrand

I worked with a large Manhattan-based newscaster who was so unable to find their own content that they had to go Google themselves to find their own web pages, just to get a clue as to who was the reporter and what the headline was, so they could go back into their content factory and try to find where it actually lived. Or go down the hall to the person who put sticky notes on tapes and hard drives and hope that helped.

Another example: a New York-based financial services company said they’re so pressed for time, because they can’t find stuff, that they literally spend millions of dollars recreating content externally rather than finding what they already have. Findability is a disaster outcome, and it’s the biggest problem most businesses, and a lot of broadcasting operations too, have to deal with. It has many causes, but it can be solved with the right solution, designed and deployed the right way.

The right solution brings intelligence as an inherent dimension of the content itself. It’s no longer just a file you have to open to see what it is. It should convey an understanding of the content in many dimensions — emotional tone, what kind of music, is this horror or action. The right solution doesn’t rely on people keying in metadata anymore. The days of hiring people to sit there feeding metadata into a reality show 24/7 are over. Other people will not enter metadata — forget about it. They might enter a file name they like, but that’s about it, unless you force them, and if you force them too much they’ll hate you.

The right solution leverages AI services, hosted in the cloud or on premises, to extract that metadata — facial recognition, scene detection, video description, transcripts, captions, mood, action — and incorporates it into the asset management system, into the content factory, so people can search by keyword, by semantic search, search it like they’d say it. And not just get results, but get suggestions — like old Clippy from Microsoft: looks like you’re working on an ad for this, may I suggest this content or that content. Many modern tools do this well individually. The challenge is bringing them together in a single platform and having it actually work. That’s the role of good advisers who understand the technology and also understand the needs of the business.

Metadata is fragile (17:25)

Jeff: This is something I’ve been talking about for months — metadata is fragile and absolutely necessary, and those two things together make this so hard. We’ve got production metadata, operational metadata, and preservation metadata. Three different layers, different lifespans, and most systems just say we’ll put it all together, then drop the two that don’t get attention. We build these solutions on the front end of the production system, and at the back end the preservation people are asking what the hell were you thinking. The operational people — that’s business metadata, where can this be used elsewhere in the business — if that’s not there, you’ve cut the legs off the lifespan of the content.

So it’s a simple test. Can you get someone in marketing, or an intern who’s never used your system, and say: I need to find the CEO talking about the New Jersey office. It was three years ago. You’ve got five minutes. Go. If they can find it without calling anyone, that’s a successful test case for how your metadata is actually working. Everyone loves the results of having metadata there. Nobody wants to put the metadata in in the first place. It’s a tiring process nobody wants to engage in.

I once spoke with a customer who wanted to define 250 metadata fields that no one was ever going to fill in.

Eivind Sandstrand

Eivind: I once spoke with a customer who wanted to define 250 metadata fields that no one was ever going to fill in.

Coming up (19:10)

Jeff: In the coming episodes, we’ll be bringing in some manufacturers and some integrators to show what we’re talking about rather than just describing it. Please hit those like and subscribe buttons. If you’ve got questions about the concepts around this, hit us up in the comments. Both of us will have our contact information popping up shortly, and we look forward to hearing from you. Thanks for joining us here today.

Eivind: Thank you very much, everybody.

Similar Posts