Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Rule of thumb is never having sensitive conversations where such exploits are feasible.


With enough processing power, you can listen to anyone as long as you can get an object in frame that’s vibrating with speech (the UK and Chinese CCTV networks comes to mind). Depending on resolution, I’d expect this to be the case for any footage already stored. It’s just a matter of having a distributed computing job kicked off to comb through video data and add the additional audio metadata.

Will you only speak of sensitive subjects in rooms with no windows? Will we be silent in public? These are issues where technology, politics, and human rights intersect.


Back when this was new, I remember discussing it on Hackaday or somewhere. Someone remarked that they were amazed that the video could pick up on the motion at all, the video looked still to them. I responded that it really was still--the demo video had been lossfully encoded for the web (re-encoded by YouTube?), and the encoder took out the subtle changes.

I suspect that most footage already stored is similarly lossfully encoded, and that this technique isn't possible on it.


>the UK and Chinese CCTV networks comes to mind

Minor nitpick, but the UK doesn't have a CCTV network. It has a huge number of privately owned CCTV cameras and a relatively small number of CCTV cameras operated by individual local authorities and police forces. The privately-owned cameras aren't joined up in any useful way and are often of very poor quality; the publicly owned cameras are overwhelmingly used for real-time monitoring of busy city centre locations.

Installing CCTV cameras is cheap and easy, but usefully monitoring them is expensive and difficult, even with whizz-bang CV algorithms. I'm deeply sceptical as to how useful any state-level CCTV network would be for mass surveillance. 20 million 4K/30fps cameras would produce something in the region of two exabytes per day; just storing that data would cost about $7bn per month.


Run the algo on any reflective objects in the field of view, feed the outputted audio to a transcription NN. Save all high confidence transcribed voice as compressed txt. Save the source video if you pass a threshold of danger keywords. ("bomb", "kill", "prime minister", "didn't pay my tv licence")


Best case scenario, you'll need a nuclear power station and a few dozen data centers full of ASICs. Oh, and an internet's worth of extra bandwidth.


"Will you only speak of sensitive subjects in rooms with no windows" - yes, the building standards for secure facilities (i.e. a room where discussing classified information above a certain grade is permitted) mandate no external windows, and much more than that.


I tried to look this up, info here: https://www.dni.gov/files/NCSC/documents/Regulations/Technic... (Ch. 3 sec. F, pg. 13)

F. SCIF Window Criteria

1. Every effort should be made to minimize or eliminate windows in the SCIF, especially on the ground floor.

2. Windows shall be non-opening.

3. Windows shall be protected by security alarms in accordance with Chapter 7 when they are within 18 feet of the ground or an accessible platform.

4. Windows shall provide visual and acoustic protection.

5. Windows shall be treated to provide RF protection when recommended by the CTTA.

6. All windows less than 18 feet above the ground or from the nearest platform affording access to the window (measured from the bottom of the window), shall be protected against forced entry and meet the standard for the perimeter.



There's also the room in a room concept where one is suspended inside the other. You might add things like ultrasound masking or blocking, too, since it was a known attack vector. From there, I thought about, but can't recall if implemented, some double doorway with buffer in between so opening a door didn't leak sounds/signals. Open one, go in, close it, and then go through other.


Don't sweat the CCTV or recordings:

Reconstructing audio from video requires that the frequency of the video samples — the number of frames of video captured per second — be higher than the frequency of the audio signal. In some of their experiments, the researchers used a high-speed camera that captured 2,000 to 6,000 frames per second. That’s much faster than the 60 frames per second possible with some smartphones, but well below the frame rates of the best commercial high-speed cameras, which can top 100,000 frames per second.


Right, which is why the researchers were reliant on video artifacts caused by the rolling shutter effect. Even at NTSC resolution, that could get you 262.5 scanlines per frame at 60 FPS (525 at ~29.75 FPS, because of interlacing... NTSC is complicated), which could potentially get you ~15 KHz as a baseline sample rate, depending on how prominent the rolling shutter effect actually is for that image sensor, and how much of the frame your vibrating object consumes vertically. This isn't perfect and would require all sorts of calibration, but I could see it being used to recover windows of time at much higher sample rates, and make inferences about the inbetween data with some FFT analysis on what you recovered.

Basically, it's more the rate of _scanlines_ that matters for this technique, and the quality of the image sensor used. The rate of full frames isn't the limiter.


The article quite literally is about a technique to bypass that limitation and do it with a regular 60fps camera.


It sounds like that technique doesn't quite recover the audio though, just a portion of it:

While this audio reconstruction wasn’t as faithful as that with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities.

Saying it may still be good enough pretty strongly implies that the quality is quite low.


There is a video on the webpage that includes the recovered audio. It is low quality, but enough to understand the speaker most of the time, and definately enough to pose a security risk when analysed by a specialist


The youtube clip in the article? The voice reconstruction there is from high speed video. If you mean some other video, I'd appreciate a link.


> Will you only speak of sensitive subjects in rooms with no windows?

That is exactly why almost all DOD secure areas are windowless.


> That is exactly why almost all DOD secure areas are windowless.

I'm pretty sure that a bigger factor in why secure facilities have limitations on windows (and especially on ground floor windows) is physical security.


Not really. Non-operator windows do not degrade physical security that much, or really at all if they are not accessible from the ground. It is primarily danger of surveillance that drives the lack of windows.


Don’t disagree. Scale that to an entire populace though.


In the brave new world that's being built, the entire population (most of it) doesn't get to have true confidentiality if somebody is sufficiently interested in them.


Well, windowless bathrooms are not uncommon. And most modern construction has walk-in closets, which typically lack windows. Even so, a source of white noise is always prudent.


And leads to mold and bad air unless you are careful which people aren't


True, but only if you don't ventilate properly. And exhaust fans are also a useful source of white noise. Along with running water.


> Will you only speak of sensitive subjects in rooms with no windows?

Of course. I mean, that's been obvious for at least a decade.

> Will we be silent in public?

There's no need to be silent. However, one must be aware of surveillance risks, and act accordingly.

> These are issues where technology, politics, and human rights intersect.

For sure. But just wanting privacy doesn't work. And you must always deal with what's so.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: