Initial commit: iPhone Live Photo clip concatenator
Detects genuine audio/video overlap between consecutive Live Photo clips via audio cross-correlation, trims the duplicated footage, and re-encodes a seamless, metadata-preserving output.
This commit is contained in:
@@ -0,0 +1,5 @@
|
||||
__pycache__/
|
||||
*.pyc
|
||||
*.mov
|
||||
*.MOV
|
||||
*.raw
|
||||
@@ -0,0 +1,80 @@
|
||||
# live-photo-concat
|
||||
|
||||
Concatenate sequential iPhone Live Photo `.MOV` clips into a single seamless,
|
||||
metadata-preserving video — with the real duplicate footage between clips
|
||||
removed.
|
||||
|
||||
## Why
|
||||
|
||||
Each iPhone Live Photo `.MOV` captures roughly 1.5 seconds before and after
|
||||
its key moment. When several Live Photos are taken in quick succession (e.g.
|
||||
rapid-fire shutter presses), consecutive clips genuinely overlap: the same
|
||||
seconds of real-world video and audio get captured twice, once at the end of
|
||||
clip *N* and again at the start of clip *N+1*. Naively concatenating the
|
||||
clips repeats that footage, which shows up as an odd "loop" at each clip
|
||||
boundary.
|
||||
|
||||
This script:
|
||||
|
||||
1. Extracts the real video+audio streams from each clip (iPhone Live Photo
|
||||
files bundle a handful of extra HDR/depth/metadata tracks that aren't
|
||||
needed here).
|
||||
2. Cross-correlates the audio at each clip boundary to measure the *actual*
|
||||
overlap duration from the content itself (not just filename order or
|
||||
whole-second creation timestamps, which aren't precise enough).
|
||||
3. Trims the duplicated span off the start of each subsequent clip.
|
||||
4. Re-encodes the video across the joins (trimming mid-GOP HEVC can't be done
|
||||
with a plain stream copy) and concatenates everything into one file.
|
||||
5. Carries over the original container metadata (GPS, device info, creation
|
||||
time, Live Photo identifiers) from one of the source clips.
|
||||
|
||||
Audio stays uncompressed PCM throughout, so it never loses quality. Video
|
||||
quality defaults to a bitrate slightly above the source clips' own bitrate
|
||||
(configurable — see below).
|
||||
|
||||
## Requirements
|
||||
|
||||
- `ffmpeg` / `ffprobe` on `PATH`
|
||||
- Python 3 with the packages in `requirements.txt`:
|
||||
|
||||
```bash
|
||||
pip install --user -r requirements.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
./concat_live_clips.py IMG_2441.MOV IMG_2442.MOV IMG_2443.MOV ...
|
||||
./concat_live_clips.py -o myvideo.mov clip1.MOV clip2.MOV ...
|
||||
```
|
||||
|
||||
If `-o/--output` is omitted, the output filename is derived from the source
|
||||
filenames' numbering, e.g. `IMG_2441.MOV .. IMG_2445.MOV` produces
|
||||
`IMG_2441-2445.mov`.
|
||||
|
||||
### Numbered sequence shortcuts
|
||||
|
||||
```bash
|
||||
# bash / zsh
|
||||
./concat_live_clips.py IMG_{2441..2445}.MOV
|
||||
|
||||
# PowerShell
|
||||
python .\concat_live_clips.py (2441..2445 | ForEach-Object { "IMG_$_.MOV" })
|
||||
```
|
||||
|
||||
### Options
|
||||
|
||||
| Flag | Description |
|
||||
|---|---|
|
||||
| `-o, --output` | Output file path (default: derived from source filenames) |
|
||||
| `--crf N` | Use CRF (quality-based) encoding instead of the default bitrate target. Lower = higher quality; ~14-18 is near-transparent |
|
||||
| `--bitrate RATE` | Explicit target video bitrate, e.g. `12M` or `12000k` |
|
||||
| `--lossless` | Mathematically lossless video encoding (much larger output) |
|
||||
| `--bitrate-multiplier N` | When auto-selecting a bitrate, multiply the source clips' peak bitrate by this (default: `1.2`) |
|
||||
| `--bitrate-floor RATE` | When auto-selecting a bitrate, never go below this (default: `10M`) |
|
||||
| `--preset` | x265 preset (default: `medium`) |
|
||||
| `--metadata-from N` | Take container metadata from the Nth input clip, 0-indexed (default: `0`, the first clip) |
|
||||
| `--confidence-threshold N` | Minimum audio cross-correlation confidence required to trust an overlap detection before falling back to 0 (default: `0.9`) |
|
||||
|
||||
`--crf`, `--bitrate`, and `--lossless` are mutually exclusive; the default
|
||||
(no flag) auto-computes a bitrate target from the source clips.
|
||||
Executable
+271
@@ -0,0 +1,271 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Concatenate sequential iPhone Live Photo .MOV clips into one seamless,
|
||||
metadata-preserving file, with real duplicate footage removed.
|
||||
|
||||
Consecutive Live Photos taken close together in time genuinely re-capture
|
||||
the same seconds of real-world video+audio (each Live Photo spans roughly
|
||||
1.5s before/after its key moment). Naively concatenating them repeats that
|
||||
footage. This script cross-correlates the audio at each clip boundary to
|
||||
measure the real overlap duration from content, then trims that duplicated
|
||||
span off the start of each subsequent clip before joining. Video is
|
||||
re-encoded (trimming mid-GOP HEVC can't be done with a plain stream copy);
|
||||
by default the target bitrate is set slightly above the source clips' own
|
||||
bitrate, so quality shouldn't visibly regress. Audio is PCM throughout, so
|
||||
it stays lossless. Metadata (GPS, device info, creation time, Live Photo
|
||||
IDs) is taken from one clip (the first, by default).
|
||||
|
||||
Usage:
|
||||
./concat_live_clips.py IMG_2441.MOV IMG_2442.MOV IMG_2443.MOV ...
|
||||
./concat_live_clips.py -o myvideo.mov clip1.MOV clip2.MOV ...
|
||||
|
||||
Numbered sequence shortcuts:
|
||||
bash/zsh: ./concat_live_clips.py IMG_{2441..2445}.MOV
|
||||
PowerShell: python .\\concat_live_clips.py (2441..2445 | ForEach-Object { "IMG_$_.MOV" })
|
||||
|
||||
Requirements:
|
||||
- ffmpeg / ffprobe on PATH
|
||||
- Python package 'numpy' (used for the audio cross-correlation that
|
||||
detects clip overlap): pip install --user numpy
|
||||
"""
|
||||
import argparse
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
try:
|
||||
import numpy as np
|
||||
except ImportError:
|
||||
sys.exit(
|
||||
"error: this script requires the 'numpy' package (used to detect "
|
||||
"audio overlap between clips), but it isn't installed.\n"
|
||||
"Install it with:\n"
|
||||
" python3 -m pip install --user numpy"
|
||||
)
|
||||
|
||||
SR = 48000
|
||||
|
||||
|
||||
def extract_audio(path, out_raw):
|
||||
subprocess.run(
|
||||
[
|
||||
"ffmpeg", "-y", "-v", "error", "-i", str(path),
|
||||
"-map", "0:a:0", "-f", "s16le", "-acodec", "pcm_s16le",
|
||||
"-ar", str(SR), "-ac", "1", str(out_raw),
|
||||
],
|
||||
check=True,
|
||||
)
|
||||
|
||||
|
||||
def load_audio(path):
|
||||
return np.fromfile(path, dtype="<i2").astype(np.float64)
|
||||
|
||||
|
||||
def best_overlap(a_tail, b_head, sr=SR, kmin_s=0.05, coarse_step_s=0.002, refine_range_s=0.01):
|
||||
"""Find the overlap (seconds) that best aligns the tail of clip A with
|
||||
the head of clip B, via normalized cross-correlation, plus a confidence
|
||||
score in [0, 1]."""
|
||||
n = min(len(a_tail), len(b_head))
|
||||
a_tail = a_tail[-n:]
|
||||
b_head = b_head[:n]
|
||||
kmin = max(1, int(kmin_s * sr))
|
||||
coarse_step = max(1, int(coarse_step_s * sr))
|
||||
|
||||
def score(k):
|
||||
a_seg = a_tail[n - k:]
|
||||
b_seg = b_head[:k]
|
||||
ea = np.dot(a_seg, a_seg)
|
||||
eb = np.dot(b_seg, b_seg)
|
||||
if ea < 1e-6 or eb < 1e-6:
|
||||
return 0.0
|
||||
return float(np.dot(a_seg, b_seg) / np.sqrt(ea * eb))
|
||||
|
||||
ks = list(range(kmin, n, coarse_step))
|
||||
if not ks:
|
||||
return 0.0, 0.0
|
||||
scores = [score(k) for k in ks]
|
||||
best_i = int(np.argmax(scores))
|
||||
best_k = ks[best_i]
|
||||
best_score = scores[best_i]
|
||||
|
||||
refine_range = max(1, int(refine_range_s * sr))
|
||||
lo = max(kmin, best_k - coarse_step - refine_range)
|
||||
hi = min(n, best_k + coarse_step + refine_range)
|
||||
for k in range(lo, hi):
|
||||
s = score(k)
|
||||
if s > best_score:
|
||||
best_score = s
|
||||
best_k = k
|
||||
|
||||
return best_k / sr, best_score
|
||||
|
||||
|
||||
def detect_overlaps(clips, confidence_threshold):
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
raws = []
|
||||
for c in clips:
|
||||
raw = Path(tmp) / (c.name + ".raw")
|
||||
extract_audio(c, raw)
|
||||
raws.append(raw)
|
||||
audios = [load_audio(r) for r in raws]
|
||||
|
||||
overlaps = []
|
||||
for i in range(len(audios) - 1):
|
||||
overlap_s, score = best_overlap(audios[i], audios[i + 1])
|
||||
if score < confidence_threshold:
|
||||
print(
|
||||
f"warning: low-confidence overlap ({score:.3f}) between "
|
||||
f"{clips[i].name} and {clips[i + 1].name}; treating as 0 overlap",
|
||||
file=sys.stderr,
|
||||
)
|
||||
overlap_s = 0.0
|
||||
print(f" {clips[i + 1].name}: trimming {overlap_s:.3f}s of duplicated "
|
||||
f"start (confidence {score:.3f})", file=sys.stderr)
|
||||
overlaps.append(overlap_s)
|
||||
return overlaps
|
||||
|
||||
|
||||
def parse_rate(rate):
|
||||
"""Parse an ffmpeg-style bitrate string ('12M', '12000k', '12000000') into bits/second."""
|
||||
rate = str(rate)
|
||||
multiplier = {"k": 1_000, "M": 1_000_000}.get(rate[-1], None)
|
||||
if multiplier:
|
||||
return int(float(rate[:-1]) * multiplier)
|
||||
return int(rate)
|
||||
|
||||
|
||||
def source_video_bitrate(clips):
|
||||
"""Peak per-clip video-stream bitrate (bits/s) across the sources, used
|
||||
as the baseline for the default output bitrate."""
|
||||
rates = []
|
||||
for c in clips:
|
||||
out = subprocess.run(
|
||||
["ffprobe", "-v", "error", "-select_streams", "v:0",
|
||||
"-show_entries", "stream=bit_rate:format=size,duration",
|
||||
"-of", "default=noprint_wrappers=1:nokey=1", str(c)],
|
||||
check=True, capture_output=True, text=True,
|
||||
).stdout.split()
|
||||
bit_rate = int(out[0]) if out and out[0] != "N/A" else None
|
||||
if bit_rate is None:
|
||||
size, duration = int(out[1]), float(out[2])
|
||||
bit_rate = int(size * 8 / duration)
|
||||
rates.append(bit_rate)
|
||||
return max(rates)
|
||||
|
||||
|
||||
NUMBERED_RE = re.compile(r"^(.*?)(\d+)(\.[^.]+)?$")
|
||||
|
||||
|
||||
def default_output_name(clips):
|
||||
"""Build an output filename from the source files' numbering, e.g.
|
||||
IMG_2441.MOV..IMG_2445.MOV -> IMG_2441-2445.mov. Falls back to
|
||||
'combined.mov' if the names don't share a recognizable numbered pattern."""
|
||||
first_match = NUMBERED_RE.match(clips[0].name)
|
||||
last_match = NUMBERED_RE.match(clips[-1].name)
|
||||
if first_match and last_match and first_match.group(1) == last_match.group(1):
|
||||
prefix = first_match.group(1)
|
||||
first_num = first_match.group(2)
|
||||
last_num = last_match.group(2)
|
||||
if first_num == last_num:
|
||||
return Path(f"{prefix}{first_num}.mov")
|
||||
return Path(f"{prefix}{first_num}-{last_num}.mov")
|
||||
return Path("combined.mov")
|
||||
|
||||
|
||||
def build_filter(clips, overlaps):
|
||||
trim_filters = ""
|
||||
concat_inputs = ""
|
||||
for i in range(len(clips)):
|
||||
if i == 0:
|
||||
concat_inputs += f"[{i}:0][{i}:6]"
|
||||
else:
|
||||
trim = overlaps[i - 1]
|
||||
trim_filters += (
|
||||
f"[{i}:0]trim=start={trim},setpts=PTS-STARTPTS[v{i}];"
|
||||
f"[{i}:6]atrim=start={trim},asetpts=PTS-STARTPTS[a{i}];"
|
||||
)
|
||||
concat_inputs += f"[v{i}][a{i}]"
|
||||
return f"{trim_filters}{concat_inputs}concat=n={len(clips)}:v=1:a=1[v][a]"
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
parser.add_argument("clips", nargs="+", type=Path, help="Live Photo .MOV clips, in playback order")
|
||||
parser.add_argument("-o", "--output", type=Path, default=None,
|
||||
help="Output file (default: derived from source filenames, e.g. IMG_2441-2445.mov)")
|
||||
quality = parser.add_mutually_exclusive_group()
|
||||
quality.add_argument("--crf", type=float, default=None,
|
||||
help="Use CRF (quality-based) encoding instead of the default bitrate target (lower = higher quality, ~14-18 is near-transparent)")
|
||||
quality.add_argument("--bitrate", type=str, default=None, metavar="RATE",
|
||||
help="Explicit target video bitrate, e.g. 12M or 12000k (default: auto, ~1.2x the source clips' own bitrate)")
|
||||
quality.add_argument("--lossless", action="store_true",
|
||||
help="Mathematically lossless encoding (much larger output)")
|
||||
parser.add_argument("--bitrate-multiplier", type=float, default=1.2,
|
||||
help="When auto-selecting a bitrate, multiply the source clips' peak bitrate by this (default: 1.2)")
|
||||
parser.add_argument("--bitrate-floor", type=str, default="10M", metavar="RATE",
|
||||
help="When auto-selecting a bitrate, never go below this (default: 10M)")
|
||||
parser.add_argument("--preset", default="medium", help="x265 preset (default: medium)")
|
||||
parser.add_argument("--metadata-from", type=int, default=0, metavar="N",
|
||||
help="Take container metadata from the Nth clip, 0-indexed (default: 0, the first clip)")
|
||||
parser.add_argument("--confidence-threshold", type=float, default=0.9,
|
||||
help="Minimum audio cross-correlation confidence to trust an overlap detection (default: 0.9)")
|
||||
args = parser.parse_args()
|
||||
|
||||
clips = args.clips
|
||||
if len(clips) < 2:
|
||||
parser.error("need at least 2 clips")
|
||||
for c in clips:
|
||||
if not c.is_file():
|
||||
parser.error(f"file not found: {c}")
|
||||
|
||||
output = args.output or default_output_name(clips)
|
||||
|
||||
print(f"Detecting real overlap between {len(clips)} clips via audio cross-correlation...", file=sys.stderr)
|
||||
overlaps = detect_overlaps(clips, args.confidence_threshold)
|
||||
|
||||
filter_complex = build_filter(clips, overlaps)
|
||||
|
||||
video_args = ["-c:v", "libx265", "-tag:v", "hvc1", "-pix_fmt", "yuv420p"]
|
||||
if args.lossless:
|
||||
video_args += ["-x265-params", "lossless=1"]
|
||||
elif args.crf is not None:
|
||||
video_args += ["-crf", str(args.crf)]
|
||||
else:
|
||||
if args.bitrate is not None:
|
||||
target = args.bitrate
|
||||
else:
|
||||
source_rate = source_video_bitrate(clips)
|
||||
floor_bps = parse_rate(args.bitrate_floor)
|
||||
target_bps = max(int(source_rate * args.bitrate_multiplier), floor_bps)
|
||||
target = f"{target_bps}"
|
||||
print(f"Source peak video bitrate ~{source_rate / 1e6:.1f} Mbps -> "
|
||||
f"targeting ~{target_bps / 1e6:.1f} Mbps"
|
||||
f"{' (bitrate floor applied)' if target_bps == floor_bps else ''}", file=sys.stderr)
|
||||
headroom = str(int(parse_rate(target) * 1.5))
|
||||
video_args += ["-b:v", target, "-maxrate", headroom, "-bufsize", headroom]
|
||||
# The source clips have irregular variable frame rates, which get worse
|
||||
# once trimmed/concatenated; that confuses x265's rate control (it badly
|
||||
# underspends the target bitrate). Normalize to a constant frame rate.
|
||||
video_args += ["-r", "30", "-preset", args.preset]
|
||||
|
||||
cmd = ["ffmpeg", "-y"]
|
||||
for c in clips:
|
||||
cmd += ["-i", str(c)]
|
||||
cmd += [
|
||||
"-filter_complex", filter_complex,
|
||||
"-map", "[v]", "-map", "[a]",
|
||||
"-map_metadata", str(args.metadata_from),
|
||||
*video_args,
|
||||
"-c:a", "pcm_s16le",
|
||||
"-movflags", "+faststart+use_metadata_tags",
|
||||
str(output),
|
||||
]
|
||||
|
||||
print(f"Encoding seamless, deduplicated output -> {output}", file=sys.stderr)
|
||||
subprocess.run(cmd, check=True)
|
||||
print(f"Done: {output}", file=sys.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1 @@
|
||||
numpy
|
||||
Reference in New Issue
Block a user