TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	R1@0.5	54.26	# 1
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	R1@0.7	38.45	# 1
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	mAP@0.5	54.17	# 1
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	mAP@0.75	29.73	# 1
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	mAP	30.91	# 1

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/vtg-gpt-tuning-free-zero-shot-video-temporal-1/zero-shot-moment-retrieval-on-qvhighlights)](https://paperswithcode.com/sota/zero-shot-moment-retrieval-on-qvhighlights?p=vtg-gpt-tuning-free-zero-shot-video-temporal-1)`

VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT

Applied Sciences 2024 · Yifang Xu, Yunzhuo Sun, Zien Xie, Benxiang Zhai, Sidan Du ·

Video temporal grounding (VTG) aims to locate specific temporal segments from an untrimmed video based on a linguistic query. Most existing VTG models are trained on extensive annotated video-text pairs, a process that not only introduces human biases from the queries but also incurs significant computational costs. To tackle these challenges, we propose VTG-GPT, a GPT-based method for zero-shot VTG without training or fine-tuning. To reduce prejudice in the original query, we employ Baichuan2 to generate debiased queries. To lessen redundant information in videos, we apply MiniGPT-v2 to transform visual content into more precise captions. Finally, we devise the proposal generator and post-processing to produce accurate segments from debiased queries and image captions. Extensive experiments demonstrate that VTG-GPT significantly outperforms SOTA methods in zero-shot settings and surpasses unsupervised approaches. More notably, it achieves competitive performance comparable to supervised methods. The code is available on https://github.com/YoucanBaby/VTG-GPT

PDF Abstract Applied Sciences 2024 PDF

Code

Add Remove Mark official

YoucanBaby/VTG-GPT official

Tasks

Add Remove

Image Captioning

Zero-shot Moment Retrieval

Datasets

ActivityNet

Charades

ActivityNet Captions

Charades-STA

QVHighlights

Results from the Paper

Add Remove

Ranked #1 on Zero-shot Moment Retrieval on QVHighlights

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Zero-shot Moment Retrieval	QVHighlights	VTG-GPT	R1@0.5	54.26	# 1	Compare
			R1@0.7	38.45	# 1	Compare
			mAP@0.5	54.17	# 1	Compare
			mAP@0.75	29.73	# 1	Compare
			mAP	30.91	# 1	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit Add Remove

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Add Remove

Methods

Add Remove