TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Video Quality Assessment	KoNViD-1k	DisCoVQA	PLCC	0.860	# 5
Video Quality Assessment	LIVE-FB LSVQ	DisCoVQA	PLCC	0.850	# 9
Video Quality Assessment	LIVE-VQC	DisCoVQA	PLCC	0.844	# 5

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/discovqa-temporal-distortion-content/video-quality-assessment-on-konvid-1k)](https://paperswithcode.com/sota/video-quality-assessment-on-konvid-1k?p=discovqa-temporal-distortion-content)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/discovqa-temporal-distortion-content/video-quality-assessment-on-live-vqc)](https://paperswithcode.com/sota/video-quality-assessment-on-live-vqc?p=discovqa-temporal-distortion-content)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/discovqa-temporal-distortion-content/video-quality-assessment-on-live-fb-lsvq)](https://paperswithcode.com/sota/video-quality-assessment-on-live-fb-lsvq?p=discovqa-temporal-distortion-content)`

DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment

20 Jun 2022 · HaoNing Wu, Chaofeng Chen, Liang Liao, Jingwen Hou, Wenxiu Sun, Qiong Yan, Weisi Lin ·

The temporal relationships between frames and their influences on video quality assessment (VQA) are still under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some temporal variations (such as shaking, flicker, and abrupt scene transitions) are causing temporal distortions and lead to extra quality degradations, while other variations (e.g. those related to meaningful happenings) do not. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design a transformer-based Spatial-Temporal Distortion Extraction (STDE) module. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. We will publish our codes and pretrained weights later.

PDF Abstract

Code

Add Remove Mark official

QualityAssessment/DisCoVQA official

Tasks

Add Remove

Time Series Analysis

Video Quality Assessment

Visual Question Answering (VQA)

Datasets

LIVE-VQC KoNViD-1k

LIVE-FB LSVQ

Results from the Paper

Edit

Ranked #5 on Video Quality Assessment on KoNViD-1k

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Video Quality Assessment	KoNViD-1k	DisCoVQA	PLCC	0.860	# 5	Compare
Video Quality Assessment	LIVE-FB LSVQ	DisCoVQA	PLCC	0.850	# 9	Compare
Video Quality Assessment	LIVE-VQC	DisCoVQA	PLCC	0.844	# 5	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove